Local LLM stack
A self-hosted vLLM, LiteLLM, Open WebUI, monitoring, and fine-tuning lab stack for an NVIDIA Spark host.
Production-shaped ML work, applied research tooling, and experiments with enough surface area to inspect.
A self-hosted vLLM, LiteLLM, Open WebUI, monitoring, and fine-tuning lab stack for an NVIDIA Spark host.
An interactive in-browser MNIST lab with ONNX Runtime inference, drawing, preprocessing traces, feature heatmaps, and linked embedding/logit spaces.
A compact PyTorch lab for fast-weight and test-time-training ideas, with toy equivalence demos and a small language-model training harness.
A summarization modeling project with data preparation, fine-tuning artifacts, and loss-curve diagnostics.
A local-first knowledge system for turning papers and notes into reviewed article nodes, concept graphs, and publishable vault output.
Random-forest-style feature bagging for high-dimensional clustering experiments.
Production-style baseline templates for classification, survival, anomaly, and time-to-event modeling.
Reliability modeling patterns for translating messy field behavior into test plans and failure-risk estimates.
A local-first ingest pipeline for transcripts, audio features, retrieval artifacts, and repeatable experiment runs.
A defect-detection lab for reconstruction, segmentation, and evaluation on industrial visual anomaly data.
Audio classification experiments around spectrograms, augmentation, validation discipline, and competition constraints.
A practical lab notebook for TabPFN, TabICL, uncertainty, conformal prediction, and tabular baselines.
A reusable repo shape for research agents: plans, artifacts, eval notes, and reproducible handoff state.
A local embedding workspace for indexing, comparing, and inspecting text representations without external services.
Audio representation experiments across waveforms, spectra, model embeddings, and retrieval-ready artifacts.
A small lab for MoE routing, load balancing, expert specialization, and failure modes.
A compact experiment around residual identity mappings, trainability, and small-model diagnostics.
Experiments around using reinforcement learning to choose what context to retain, compress, or drop.
A smaller reinforcement-learning lab for compaction policies, rewards, and evaluation traces.
Small reinforcement-learning environments and experiments grounded in Sutton-style examples.
A demo of online training traces, incremental metrics, and model behavior while data arrives.
A local-first pipeline for audio download, Whisper transcription, text embeddings, and audio embeddings.
Program evolution experiments that track fitness, complexity, and the tradeoff between improvement and bloat.
An interactive app for concentration phenomena, tails, and the geometry of probability bounds.
An interactive visualization app for eigenvalue clouds, spectra, and random-matrix intuition.
An interactive systems app for stepping through consensus messages, timing, and agreement behavior.
A browser playground for Tierra-style program evolution, mutation, replication, and population dynamics.
An interactive evolutionary-art playground for selection, mutation, and visual search.
A small, explicit experiment harness for SFT/DPO and inference-time scaling (best-of-N, verifiers).
A UI experiment around energy, inference, and controllable visual exploration.
A reproduction and extension of program-soup ALife experiments: BFF tapes, replicators, and phase-transition-like dynamics.
A recommender-system project around event affinity, user behavior signals, and retrieval-ready recommendations.
An applied modeling project for finding high-residual prescribing patterns after accounting for expected variation.
Credible intervals and Bayesian regression for deciding which marketing channels are signal versus noise.
An end-to-end lead scoring project across data cleaning, feature engineering, model validation, and deployment shape.