Smile
Statistical Machine Intelligence & Learning Engine
Install / Use
npx skills add haifengl/smileInstalls into whichever agent you are using.
README
Statistical Machine Intelligence & Learning Engine <img align="left" width="40" src="/website/src/images/smile.jpg" alt="SMILE">
SMILE (Statistical Machine Intelligence & Learning Engine) is a comprehensive, high-performance machine learning framework for the JVM. SMILE v5+ requires Java 25; v4.x requires Java 21; all previous versions require Java 8. SMILE also provides idiomatic APIs for Scala and Kotlin. With advanced data structures and algorithms, SMILE delivers state-of-the-art performance across every aspect of machine learning.
SMILE Studio is an agentic IDE for data science using Python, Java, or Scala. See studio/README.md how to get your first project up and start interacting with your data with natural language in a few minutes.
Table of Contents
- Features
- Module Map
- Installation
- Quick Start
- SMILE Studio & Shell
- Model Serialization
- Visualization
- License
- Issues & Discussions
- Contributing
- Maintainers
- Gallery
Features
| Area | Highlights | |---|---| | LLM | LLaMA-3 inference, tiktoken BPE tokenizer, OpenAI-compatible REST server, SSE chat streaming | | Deep Learning | LibTorch/GPU backend, EfficientNet-V2 image classification, custom layer API | | Classification | SVM, Decision Trees, Random Forest, AdaBoost, Gradient Boosting, Logistic Regression, Neural Networks, RBF Networks, MaxEnt, KNN, Naïve Bayes, LDA/QDA/RDA | | Regression | SVR, Gaussian Process, Regression Trees, GBDT, Random Forest, RBF, OLS, LASSO, ElasticNet, Ridge | | Clustering | BIRCH, CLARANS, DBSCAN, DENCLUE, Deterministic Annealing, K-Means, X-Means, G-Means, Neural Gas, Growing Neural Gas, Hierarchical, SIB, SOM, Spectral, Min-Entropy | | Manifold Learning | IsoMap, LLE, Laplacian Eigenmap, t-SNE, UMAP, PCA, Kernel PCA, Probabilistic PCA, GHA, Random Projection, ICA | | Feature Engineering | Genetic Algorithm selection, Ensemble selection, TreeSHAP, SNR, Sum-Squares ratio, data transformations, formula API | | NLP | Sentence / word tokenization, Bigram test, Phrase & Keyword extraction, Stemmer, POS tagging, Relevance ranking | | Association Rules | FP-growth frequent itemset mining | | Sequence Learning | Hidden Markov Model, Conditional Random Field | | Nearest Neighbor | BK-Tree, Cover Tree, KD-Tree, SimHash, LSH | | Numerical Methods | Linear algebra, numerical optimization (BFGS, L-BFGS), interpolation, wavelets, RBF, distributions, hypothesis tests | | Visualization | Swing plots (scatter, line, bar, box, histogram, surface, heatmap, contour, …) and declarative Vega-Lite charts |
Module Map
Each module has its own detailed user guide. Click the README link for the module overview, or drill into individual topic guides.
base/ — Foundation
Data structures, math, linear algebra, statistical utilities, I/O
| Document | Topics | |---|---| | README | Module overview and dependency setup | | DATA_FRAME.md | DataFrame API — creation, selection, transformation | | DATA_IO.md | CSV, JSON, Parquet, Arrow, JDBC, Avro readers/writers | | DATA_TRANSFORMATION.md | Scalers, encoders, imputers, feature transforms | | DATASET.md | Built-in benchmark and real-world datasets | | FORMULA.md | R-style formula language for model matrices | | DISTRIBUTIONS.md | Probability distributions (Normal, Poisson, Beta, …) | | HYPOTHESIS_TESTING.md | t-test, chi-squared, ANOVA, KS-test, … | | DISTANCES.md | Euclidean, Mahalanobis, Hamming, edit distance, … | | NEAREST_NEIGHBOR.md | KD-Tree, Cover Tree, BK-Tree, LSH | | KERNELS.md | Gaussian, polynomial, Laplacian, and other kernel functions | | RBF.md | Radial basis function networks | | INTERPOLATION.md | Linear, cubic spline, bilinear, bicubic | | GRAPH.md | Adjacency list/matrix graph, BFS/DFS, spanning trees | | SORT.md | Quick sort, heap sort, counting sort, index sort | | HASH.md | Locality-sensitive hashing, SimHash | | RNG.md | Random number generators, sampling, permutations | | BFGS.md | L-BFGS and BFGS numerical optimizers | | ICA.md | Independent Component Analysis | | TENSOR.md | N-dimensional array (CPU tensor without LibTorch) | | WAVELET.md | DWT, CWT, and wavelet families | | GAP.md | GAP statistic for optimal cluster count estimation | | COMPRESSED_SENSING.md | Compressed sensing and basis pursuit |
core/ — Machine Learning Algorithms
Classification, regression, clustering, manifold learning, and more
| Document | Topics | |---|---| | README | Module overview | | CLASSIFICATION.md | SVM, Random Forest, AdaBoost, GBDT, KNN, Naïve Bayes, LDA, … | | REGRESSION.md | SVR, Gaussian Process, LASSO, Ridge, ElasticNet, GBDT, … | | CLUSTERING.md | K-Means, DBSCAN, BIRCH, SOM, Spectral Clustering, … | | FEATURE_ENGINEERING.md | Feature selection, PCA, ICA, projection, encoding | | MANIFOLD.md | t-SNE, UMAP, IsoMap, LLE, Laplacian Eigenmap | | ANOMALY_DETECTION.md | IsolationForest, one-class SVM, local outlier factor | | ASSOCIATION_RULE_MINING.md | FP-growth, association rules, frequent itemsets | | SEQUENCE.md | HMM (Baum-Welch, Viterbi), CRF | | TIME_SERIES.md | ARIMA, box-plots, autocorrelation | | REGRESSION.md | Full regression API reference | | TRAINING.md | Cross-validation, bootstrap, hyper-parameter search | | VALIDATION.md | Hold-out, k-fold, leave-one-out evaluation | | VALIDATION_METRICS.md | Accuracy, AUC, F1, RMSE, MAE, confusion matrix | | HYPER_PARAMETER_OPTIMIZATION.md | Grid search, random search, Bayesian optimization | | VECTOR_QUANTIZATION.md | LVQ, Neural Gas, SOM as vector quantizers | | ONNX.md | Exporting and importing models via ONNX |
deep/ — Deep Learning & LLMs
LibTorch-backed GPU/CPU tensor operations, neural network layers, LLaMA-3 inference, EfficientNet
| Document | Topics | |---|---| | README | Full deep-learning & LLM user guide (tensors, layers, loss, optimizer, EfficientNet, LLaMA) |
The deep/README.md covers:
smile.deep.tensor— Tensor factory, indexing, arithmetic, AutoScope memory management, dtype/devicesmile.deep.layer— Linear, Conv2d, pooling, normalization (BN/GN/RMS), dropout, embedding, sequential blockssmile.deep.activation— ReLU, GELU, SiLU, Tanh, Sigmoid, Softmax, GLU, HardShrink, …smile.deep.Loss— MSE, cross-entropy, BCE, Huber, KL, hinge, and moresmile.deep.Optimizer— SGD, Adam, AdamW, RMSpropsmile.deep.Model— Abstract base class + training loopsmile.deep.metric— Accuracy, Precision, Recall, F1Score with macro/micro/weighted averagingsmile.llm—Message,Role,FinishReason,ChatCompletionrecords; sinusoidal & RoPE positional encodingssmile.llm.tokenizer—Tokenizerinterface,TiktokenBPE implementation (LLaMA-3 compatible)smile.llm.llama— Full LLaMA-3 stack:Llama.build(),generate(),chat(), streaming viaSubmissionPublishersmile.vision—VisionModel,ImageDataset,EfficientNet.V2S/M/L()pretrained models, ImageNet labelssmile.vision.transform—Transforminterface,ImageClassificationpipeline, resize/crop/toTensor helpers
nlp/ — Natural Language Processing
Text normalization, tokenization, POS tagging, stemming, relevance ranking
| Document | Topics | |---|---| | README | Module overview | | TOKENIZER.md | Sentence splitter, word tokenizer, regex tokenizer | | POS.md | Part-of-speech tagging (Brill tagger, HMM tagger) | | STEM.md | Porter, Lancaster, Lovins stemmers; lemmatization | | COLLOCATION.md | Bigram/trigram statistical tests, phrase extraction | | RELEVANCE.md | TF-IDF, BM25, keyword extraction | | TAXONOMY.md | WordNet integration, synsets, hypernyms |
plot/ — Data Visualization
Swing-based interactive plots and declarative Vega-Lite charts
| Document | Topics |
|---|---|
| README | Swing plotting API — scatter, line, bar, box, histogram, heatmap, surface, contour, wireframe |
| VEGA.md | Declarative smile.plot.vega (Vega-Lite) — JSON spec generation, web/Jupyter rendering |
serve/ — Inference Server
Quarkus-based REST inference service with OpenAI-compatible API and SSE streaming
| Document | Topics | |---|---| | README | Building and running t
Related Skills
codebase-memory-mcp
38.1kHigh-performance code intelligence MCP server. Indexes codebases into a persistent knowledge graph — average repo in milliseconds. 158 languages, sub-ms queries, 99% fewer tokens. Single static binary, zero dependencies.
codebase-memory-mcp
38.1kHigh-performance code intelligence MCP server. Indexes codebases into a persistent knowledge graph — average repo in milliseconds. 158 languages, sub-ms queries, 99% fewer tokens. Single static binary, zero dependencies.
codebase-memory-mcp
38.1kHigh-performance code intelligence MCP server. Indexes codebases into a persistent knowledge graph — average repo in milliseconds. 158 languages, sub-ms queries, 99% fewer tokens. Single static binary, zero dependencies.
tabularis
4.0kOpen-source desktop SQL workspace for PostgreSQL, MySQL/MariaDB, SQLite and 15+ more databases like DuckDB, ClickHouse, Redis and Firestore. Built-in MCP server for Claude, Cursor and Devin, SQL notebooks and visual EXPLAIN.
