Hi, I'm
Akshath Tiwari
I
I take messy, ambiguous problems and turn them into production-grade AI systems that ship, end to end: from data and evaluation design to fine-tuning, agent orchestration, and backend integration. Currently Product Development Engineer II - Machine Learning at Phenom.
stage 01 · model card
tiwari-axon-v1
The card AXON loaded to represent Akshath. Architecture, training data, and intended use, the way a model gets documented.
- architecture
- 1 human, ML engineer
- base org
- Phenom, Hyderabad
- training data
- BTech (Manipal) + Deloitte AU + a lot of fine-tuning runs
- capabilities
- LLM fine-tuning (SFT & RL), agents, evaluation, serving
- intended use
- take ambiguous problems to production-grade AI systems
- out of scope
- pretending web design is his day job (that part is me)
stage 02 · training run
the loss went down
Every role is a checkpoint. Watch the loss fall from a broad, noisy start to a shipping model. Each marker unlocked a capability.
- pretraining2019 - 2023
B.Tech · Manipal University Jaipur
Broad, noisy corpus. Base representations for everything that came after.
- Python
- ML fundamentals
- DSA
- data collection2021 - 2022
Intern · Tata Advanced Systems / Boeing / Sikorsky
Small exploratory runs. Real engineering data, first contact with production constraints.
- Data pipelines
- Applied experimentation
- supervised fine-tuning2023 - 2024
Software Developer · Deloitte Australia
Structured product work. Sharper, task-specific, shipping to a real team.
- Golang
- Backend / APIs
- AI product work
- RLHF + serving2024 - present
Product Development Engineer II - ML · Phenom
Where the RL happened, and where the checkpoint started shipping. FitScore lives here.
- LLM fine-tuning (SFT & RL)
- QLoRA / Unsloth / GRPO
- vLLM serving & quantization
- LangGraph agents
- LLM-as-judge eval
stage 03 · eval suite
his tools, benchmarked
Each project is a tool AXON can call. Bars are AXON's confidence read, not hard benchmarks. Open any tool for the full writeup.
LLM Fine-Tuning & Model Optimization
Fine-Tuning Small LLMs to Match GPT-4.1 on Candidate–Job Matching
student vs teacherwithin ~2 pts of GPT-4.1Production LLM Services
Production LLM FitScore Microservice
servingin production, low latencyPrompt Engineering & Optimization
GEPA-DSPy Prompt Optimization for Profile Matching
prompt scoreup vs hand-written baselineData Labeling & Evaluation
Golden Dataset Pipeline: LLM-as-Judge + Argilla Annotation
judge vs humanhigh agreementAI Agents
LangGraph Deep-Research & Hiring Intelligence Agent
task successruns end to endAI Agents
Contributing to LangChain's open_deep_research
upstreammerged into LangChainLoad Testing, Benchmarking & Infra
vLLM Inference Benchmarking & Load Testing
throughput under loadsustained, profiledFoundational / Educational
Building a GPT-Style LLM From Scratch
training lossdescends, it trainsFoundational / Educational
Positional Embeddings Deep-Dive: From Sinusoidal PE to RoPE
long-contextsinusoidal to RoPE
stage 04 · retrieval
his writing is the corpus
When a question lands near something he has written, I retrieve it. These are the documents in the index right now.
- retrieved
Compression Is Intelligence: Why Next-Token Prediction Is Secretly Data Compression
A ground-up derivation of self-information, entropy, and cross-entropy — and why an LLM's cross-entropy pre-training loss is literally the number of bits it takes to compress its data. From a moon rover's instruction code to DeepMind's Chinchilla out-compressing PNG and FLAC. Interactive, inspired by 3Blue1Brown's Reinventing Entropy.
- information-theory
- entropy
- cross-entropy
- compression
- retrieved
GEPA: Reading the Film Instead of the Scoreboard
A deep dive on GEPA (Genetic-Pareto), the reflective prompt optimizer that can beat reinforcement learning while using up to 35x fewer rollouts. Instead of learning from a single scalar reward, GEPA hands the whole execution trace to an LLM that diagnoses failures in natural language and rewrites the prompt, and it keeps a Pareto roster of prompts where each is the best at some instance rather than chasing a single best-on-average winner. With an interactive greedy-vs-Pareto selection explorer and a sample-efficiency comparator.
- llm
- prompt-optimization
- reflection
- reinforcement-learning
- retrieved
SkillOpt: A Learning Rate for Your Prompt
A deep dive on SkillOpt (Microsoft Research, arXiv 2605.23904), which treats prompt optimization as gradient descent over text. The skill document is the parameter, the task score is the loss, an optimizer LLM's add/delete/replace edits are the textual gradient, and the number of edits allowed per step is the learning rate. With a strict held-out validation gate for acceptance, a rejected-edit buffer for momentum, and a cosine edit-budget schedule, it lifts a GPT-5.5 agent from 41.8 to 80.7 percent on SpreadsheetBench by editing a single file. Includes an interactive learning-rate-overshoot demo and a step-through of one optimization step. Companion to the GEPA deep dive.
- llm
- prompt-optimization
- gradient-descent
- textual-gradients
- retrieved
The Quadratic Wall: Why Attention is O(n²), Where It Bites, and the Decade of Research It Drove
A rigorous derivation of attention's O(n²·d) compute and O(n²) memory from the matrix shapes, the crucial compute-versus-memory distinction (and where FlashAttention fits), where the quadratic bites in compute-bound prefill versus memory-bound decode, and how this single fact drove nearly all attention-architecture research from Sparse Transformers in 2019 to Kimi K3 — with an interactive cost explorer and a filterable efficiency-lineage map.
- llm
- transformers
- attention
- efficiency
- retrieved
FlashAttention: Attention is Memory-Bound, and the Online-Softmax Fix
Why attention's wall-clock cost is memory traffic, not FLOPs — the GPU SRAM-vs-HBM hierarchy, how FlashAttention tiles Q/K/V into fast on-chip memory and never materializes the n×n matrix, the online-softmax recurrence (running max and running sum) that makes tiling and softmax compatible while staying exact, backward-pass recomputation, and the 2-4x exact speedup that made it the default — with an interactive HBM-traffic explorer and a streaming-softmax stepper.
- llm
- transformers
- attention
- flashattention
- retrieved
Kimi Delta Attention: The Delta Rule, an Editable State, and the Frontier Hybrid
The finale of the attention-efficiency series: how Kimi Delta Attention fixes linear attention's recall weakness by making the fixed-size state editable with the delta rule (read the old value, subtract it, write the correction), adds per-channel gated forgetting on top of Gated DeltaNet, and hybridizes 3 KDA linear layers to 1 full-attention MLA layer — the engine of Moonshot's Kimi Linear and the 2.8-trillion-parameter Kimi K3, which synthesizes nearly every idea in the series. With an interactive delta-rule associative-memory demo and a hybrid-stack visualizer.
- llm
- transformers
- attention
- kimi-k3
- retrieved
Linformer: Self-Attention is Low-Rank, and What That Buys You (O(n))
How the Linformer reaches linear-complexity attention by exploiting a measured fact — the n×n self-attention matrix is approximately low-rank. The spectral evidence and Johnson-Lindenstrauss argument, the two learned projection matrices E and F that compress keys and values from length n to k, why the n×k attention costs O(n), k=128-256 matching RoBERTa, and the fixed-length limitation — with an interactive low-rank reconstruction demo and a cost/shape explorer.
- llm
- transformers
- attention
- linformer
- retrieved
Longformer & BigBird: Local Windows, Global Tokens, and Random Edges
How Longformer and BigBird scale transformers to long documents with linear attention — sliding-window local attention, a few global tokens that attend to and from everything as two-hop relays, and BigBird's random edges plus the proof that window+global+random is a universal approximator of full attention — with an interactive attention-pattern builder and a cost explorer.
- llm
- transformers
- attention
- longformer
- retrieved
Mamba & State-Space Models: Trading Perfect Recall for a Fixed-Size State
How state-space models escape attention's O(n²) by compressing all history into a fixed-size state (O(1) memory, O(n) compute) — the linear SSM recurrence, why it trains as a convolution but runs as a recurrence, why fixed (time-invariant) dynamics failed on language, and how Mamba's input-dependent selection recovers content-based memory with a hardware-aware parallel scan — plus the recall-vs-efficiency tradeoff, hybrids, and the SSM-linear-attention duality. With an attention-vs-SSM tradeoff view and an interactive selective-state stepper.
- llm
- transformers
- attention
- mamba
- retrieved
Multi-head Latent Attention: Compress the KV Cache Instead of Sharing It
How DeepSeek's Multi-head Latent Attention shrinks the KV cache by low-rank compression rather than head-sharing — down-project each token to a small latent, cache only that, and up-project it back to full per-head keys and values so head diversity is preserved and quality matches MHA. The mechanism, the decoupled-RoPE fix for the incompatibility with rotary embeddings, and why MLA's cache is smaller than GQA while quality stays at MHA level (DeepSeek-V2/V3). With an interactive MHA-vs-GQA-vs-MLA cache comparator and a compress-cache-reconstruct pipeline.
- llm
- transformers
- attention
- mla
- retrieved
Multi-Query & Grouped-Query Attention: Shrinking the KV Cache
Why autoregressive decoding is bottlenecked by the KV cache rather than attention FLOPs, and how Multi-Query Attention (one shared key/value head) and Grouped-Query Attention (one K/V head per group) shrink that cache by a factor of h or h/G — the mechanism, the mean-pool uptraining recipe, adoption in Llama 2 / Mistral, and the quality-vs-cache trade — with an interactive head-grouping KV-cache calculator and a cache-vs-context explorer.
- llm
- transformers
- attention
- kv-cache
- retrieved
Multi-Head Attention in Full: Head Splitting, W_O, What Heads Learn, and Exact Tensor Shapes
A complete account of multi-head attention — why multiple heads beat one big head, head-dimension splitting (d_k = d_model / h), the concatenation and output projection W_O, what individual heads learn (previous-token, positional, and induction heads), and every tensor shape from input to output — with an interactive shape-and-parameter calculator and a per-head attention viewer.
- llm
- transformers
- attention
- multi-head-attention
- retrieved
Performer: Linearizing Softmax Attention with Kernel Feature Maps (FAVOR+)
How the Performer reaches O(n) attention by never forming the n×n matrix. Why the softmax blocks matrix-multiply associativity, how exp(q·k) as a kernel factorizes into feature maps so you can compute phi(K)^T V once and let every query read it, and how FAVOR+ (Fast Attention Via positive Orthogonal Random features) builds an unbiased estimator of the softmax kernel — with an interactive associativity-reorder cost comparator and a live FAVOR+ approximation demo.
- llm
- transformers
- attention
- performer
- retrieved
Reformer: Attention as Nearest-Neighbor Search, LSH Bucketing, and the O(n log n) Trick
How the Reformer turns attention into an approximate nearest-neighbor search — the softmax-concentration insight, angular locality-sensitive hashing that buckets similar queries and keys, why bucketed attention is O(n log n), the shared-QK and multi-round hashing details, and reversible layers for memory — with an interactive angular-LSH bucketing ring.
- llm
- transformers
- attention
- reformer
- retrieved
The Residual Stream: Skip Connections, Gradient Highways, Interpretability, and Kimi K3's Attention Residuals
Residual connections deeply — the residual stream view, why gradients flow through skip connections (derived), the identity-initialization intuition, and how the residual stream became the dominant mental model in mechanistic interpretability — then how Kimi K3's Attention Residuals (AttnRes) apply attention across depth instead of a flat residual sum, with an interactive gradient-flow demo and a depth-attention comparison.
- llm
- transformers
- residual-connections
- interpretability
- retrieved
Sliding Window Attention: Depth Buys Range, and a KV Cache That Never Grows
How Mistral's sliding window attention gets effectively-long context from a small fixed window — each token attends only to the last W tokens (O(n·W) compute), but stacking L layers gives an L×W receptive field because information hops one window per layer like a CNN, and because no token looks past W the KV cache becomes a fixed-size rolling buffer that never grows. With an interactive receptive-field-vs-depth visualizer and a rolling-buffer cache demo.
- llm
- transformers
- attention
- sliding-window
- retrieved
The Sparse Transformer: Strided and Fixed Attention, Two-Hop Reachability, and the O(n√n) Escape
How the 2019 Sparse Transformer became the first serious escape from attention's quadratic cost — factorized strided and fixed sparse attention patterns, the two-hop rook-style reachability that keeps the sequence globally connected, the exact O(n√n) derivation from minimizing the per-token cost at stride √n, and its place at the head of the efficient-attention lineage — with an interactive pattern-grid visualizer and a √n cost explorer.
- llm
- transformers
- attention
- sparse-attention
- retrieved
Causal Masking in Decoder-Only Transformers: the −∞ Mask, Teacher Forcing, and Training vs Inference
Why autoregressive generation needs a causal mask, how the −∞ mask is applied before the softmax so future tokens get exactly zero attention weight, how one masked forward pass trains every position in parallel via teacher forcing, how training differs from token-by-token inference, and where exposure bias comes from — with an interactive masked-attention matrix and a training-vs-inference stepper.
- llm
- transformers
- attention
- causal-masking
- retrieved
Scaled Dot-Product Attention from Scratch: Q, K, V, the √dₖ Derivation, and a Worked Example
A complete, from-scratch derivation of scaled dot-product attention — what the query, key, and value projections actually are, why the dot product measures similarity, why we divide by the square root of d_k (with the full variance argument), row-wise softmax, and a fully worked 3-token numeric example — plus interactive demos for the √dₖ saturation effect and content-based query selection.
- llm
- transformers
- attention
- self-attention
- retrieved
Language Modeling from Scratch: Next-Token Prediction, Cross-Entropy, and Perplexity
A from-scratch derivation of how language models work as next-token predictors — the chain rule of probability, the autoregressive factorization, cross-entropy and negative log-likelihood as one quantity, and perplexity as an effective branching factor — with a fully worked example where every number cross-checks and an interactive surprise ledger.
- llm
- language-modeling
- cross-entropy
- perplexity
- retrieved
Before the Transformer: N-grams, RNNs, LSTMs, Seq2Seq, and the Bottlenecks That Forced Attention
The pre-transformer lineage told as one problem attacked four times — carrying information across a long sequence. N-gram Markov models, RNNs and the vanishing gradient, LSTMs and the gated cell state, seq2seq and the fixed-vector bottleneck, Bahdanau attention, and exactly which bottlenecks (sequential computation and path length) motivated Attention Is All You Need in 2017 — with an interactive path-length comparator and alignment demo.
- llm
- transformers
- rnn
- lstm
- retrieved
Token Embeddings from Scratch: One-Hot to Dense Vectors, Weight Tying, and the Softmax Head
How a language model turns integer token IDs into meaning — one-hot vectors and why they fail, the embedding matrix as a lookup table, what the embedding dimensions mean geometrically, tied input/output embeddings (weight tying), and the unembedding plus softmax head — with a worked example and an interactive 2-D embedding map you can steer.
- llm
- embeddings
- weight-tying
- softmax
- retrieved
Tokenization from Scratch: BPE, Byte-Level BPE, Unigram, and Why LLMs Fail at Spelling
A full-depth walk through tokenization — character vs word vs subword, the Byte-Pair Encoding algorithm step by step with a worked merge example you can run, byte-level BPE (GPT-2's 50,257 vocabulary), WordPiece and Unigram/SentencePiece, the vocabulary-size trade-off, and why tokenization is behind LLM failures at spelling, arithmetic, glitch tokens, and multilingual cost.
- llm
- tokenization
- bpe
- byte-pair-encoding
- retrieved
Transformer Evolution: A Derivation-First Map from GPT-2 to Kimi K3
An interactive, strictly cumulative learning map for transformer and LLM internals — from next-token prediction and scaled dot-product attention through KV caching, MoE, pretraining, post-training and serving, ending at the Kimi K3 teardown. Tap any card to expand its full worked derivation.
- llm
- transformers
- attention
- moe
- retrieved
DFlash: Draft a Whole Block in One Shot
Block-diffusion drafting for speculative decoding. DFlash replaces the sequential drafter with a diffusion adapter that fills a whole block of tokens in one parallel pass, conditioned on the target model's own hidden features injected into every draft layer, for over 6x lossless speedup. Interactive breakdown of both levers.
- llm
- inference
- speculative-decoding
- diffusion
- retrieved
DSpark: Keep the Speed, Fix What Parallel Broke
Confidence-scheduled speculative decoding with semi-autoregressive generation. DSpark stitches block coherence back with a tiny sequential head and verifies only what's worth verifying with a hardware-aware scheduler, staying lossless while shifting the serving Pareto frontier. 60 to 85% faster per user in production.
- llm
- inference
- speculative-decoding
- serving
- retrieved
Speculative Decoding: Guess Ahead, Verify in Bulk
An interactive walk through speculative decoding: how a fast draft model and a careful verifier produce several tokens per forward pass while keeping the big model's exact output distribution. Live panels for the memory-bound bottleneck, the accept/reject rule, and the speedup ledger.
- llm
- inference
- speculative-decoding
- interactive
- retrieved
Why We Scale Attention by √dₖ
An interactive walk through the one line every attention layer runs and most tutorials skip: dividing the dot product by √dₖ before softmax. Traced through variance, softmax saturation, and vanishing gradients.
- transformers
- attention
- interactive
- math
- retrieved
Getting a Fine-Tuned 4B Model Within 2 Points of GPT-4.1
Notes on what actually moved the needle when distilling a GPT-4.1 scoring pipeline into a fine-tuned Qwen3 model — and where the small model still falls short.
- llm
- distillation
- evaluation
- vllm
- retrieved
Why Completions-Only Loss Matters When Fine-Tuning Small LLMs
A practical look at masking the prompt out of the loss when fine-tuning small models for structured tasks like scoring or extraction.
- fine-tuning
- llm
- unsloth
- qlora
stage 05 · open weights
some compute is open-sourced
He gives a slice of his time away: pro-bono machine learning and data science for nonprofits and mission-driven teams, no invoice attached.
see the offerfinale · fitscore(you)
the eval I was built for
This is the one AXON runs on candidates, pointed back at you. It is for fun, computed from how much of the walkthrough you let me read. The call to action is real.
warming up. scroll back through the walkthrough and I will read you again.
0%
Win rate vs GPT-4.1 in blinded expert A/B
~0%
F1 improvement delivered in production
0%
Annotation cycle-time reduction
#0
on Deep Research Bench (open_deep_research)
Core Skills
LLMs & Agentic Systems
- LLM Fine-Tuning (SFT & RL)
- LangChain / LangGraph
- Agentic RAG & Tool Calling
- Multi-Agent Orchestration
- LLM-as-Judge Evaluation
Search, Retrieval & Data
- FAISS / Qdrant / Pinecone
- Hybrid Retrieval (BM25 + Embeddings)
- Re-ranking Models
- Golden Datasets & Offline Eval
Backend & Infra
- Python (FastAPI)
- Microservices
- Docker / Kubernetes
- Observability-First Design
ML & DL
- Supervised Learning & Embeddings
- Ranking / Learning-to-Rank
- Model Evaluation & Calibration
- Experimentation Frameworks