GPT-6 Astra, Looped Transformers, and Hidden ReasoningA Look at Recurrent Depth, Hidden Chains of Thought, and Recent Research on Looping Transformer Blocks09-0941268
How Claude Watermarks AI-Generated TextA 48-Minute Video Walkthrough of Token Sampling, Watermark Detection, and Removal08-2216124
Building an AI Text Detector From ScratchAn End-to-End Project With Dataset Construction, Model Training, Local Deployment, and RLVR08-15476
Controlling Reasoning Effort in LLMsHow LLMs Learn Low-, Medium-, and High-Effort Reasoning Modes07-1839745
Using Local Coding AgentsUsing Open-Weight Models in Local Coding Harnesses as an Alternative to Claude Code and Codex Subscriptions06-2746950
LLM Research Papers: The 2026 List (January to May)A curated roundup of notable LLM research papers that came out this year06-06933
Recent Developments in LLM Architectures: KV Sharing, mHC, and Compressed AttentionFrom Gemma 4 to DeepSeek V4, How New Open-Weight LLMs Are Reducing Long-Context Costs05-1634920
My Workflow for Understanding LLM ArchitecturesA Learning-Oriented Workflow for Understanding New Open-Weight Model Releases04-18824
Components of A Coding AgentHow Coding Agents Use Tools, Memory, and Repo Context to Make LLMs Work Better in Practice04-0497565
A Visual Guide to Attention Variants in Modern LLMsFrom MHA and GQA to MLA, Sparse Attention, and Hybrid Architectures03-2245916
A Dream of Spring for Open-Weight LLMs: 10 Architectures from Jan-Feb 2026A Round Up And Comparison of 10 Open-Weight LLM Releases in Spring 202602-2522112
Categories of Inference-Time Scaling for Improved LLM ReasoningAnd an Overview of Recent Inference-Scaling Papers (Including Recursive Language Models)01-24490
The State Of LLMs 2025: Progress, Problems, and PredictionsA 2025 review of large language models, from DeepSeek R1 and RLVR to inference-time scaling, benchmarks, architectures, and predictions for 2026.2025-12-3053237
LLM Research Papers: The 2025 List (July to December)In June, I shared a bonus article with my curated and bookmarked research paper lists to the paid subscribers who make this Substack possible.2025-12-30390
From DeepSeek V3 to V3.2: Architecture, Sparse Attention, and RL UpdatesUnderstanding How DeepSeek's Flagship Open-Weight Models Evolved2025-12-0328715
Beyond Standard LLMsLinear Attention Hybrids, Text Diffusion, Code World Models, and Small Recursive Transformers2025-11-0439028
Understanding the 4 Main Approaches to LLM Evaluation (From Scratch)Multiple-Choice Benchmarks, Verifiers, Leaderboards, and LLM Judges with Code Examples2025-10-0539029
Understanding and Implementing Qwen3 From ScratchA Detailed Look at One of the Leading Open-Source LLMs2025-09-061348
From GPT-2 to gpt-oss: Analyzing the Architectural AdvancesAnd How They Stack Up Against Qwen32025-08-0964248
The Big LLM Architecture ComparisonFrom DeepSeek V3 to GLM-5: A Look At Modern LLM Architecture Design2025-07-192053102