Rudrite Explainers — any idea, made visible
Visual explainers spanning high-level system design down to line-by-line code walkthroughs — every idea redrawn, animated, and made legible. Free and open.
- The System Design Cheat Sheet
- The Audit-Proof Ledger
- Design a Trading Dashboard
- Design a Chat System
- Design a Video Platform
- Design a Payment System
- Design an Email Service
- Design a Notification System
- Build a Report and Fan It Out to Millions
- Design a Proximity Service
- Top-K & Heavy Hitters
- Design an Embedding Retrieval System
- Design a Transaction Log on Object Storage
- Design a Matching Engine
- Design a Sandboxed Code-Execution Service
- Evals & Experimental Design
- Optimization Dynamics: Why Adam, Why Warmup, Why Cosine
- Mixture of Experts, Routed Honestly
- Design a Recommendation System (in the LLM Era)
- How to Design an ML System in 45 Minutes
- Design a Ranking System
- ML Reliability in Production
- GRPO Advantage: Z-Score Your Siblings, Line by Line
- GPU Arithmetic: the Numbers That Decide ML Systems
- What "Atomic" Means on a Filesystem vs. an Object Store
- Zero-RPC Sharding: How 1,000 Hosts Agree Who Writes Which Bytes
- The B-Tree Merge: Thousands of Checkpoint Files into One Atomic Manifest
- Quantization for Deployment
- How Qwix Quantizes Any Flax Model Without Touching Its Code
- GPTQ, Line by Line: Hessian-Based Weight Quantization
- The Post-Training Pipeline (SFT → RLHF → DPO)
- Scaling Laws, Honestly
- Diffusion Models, Honestly
- Design a RAG System
- Tracing → jaxpr: the One Trick Behind Every JAX Transform
- Sharding in JAX
- The SPMD Partitioner: One Program Across a Device Mesh
- The Checkpoint Lifecycle: What Happens Between save() and Durable
- Anatomy of a FlashAttention Kernel
- Inside XLA: ~200 Passes and the Fusion Decision
- How Google Spanner Works
- Design an LLM Serving Platform
- Design a Text-to-Image Service
- The Transformer, End to End
- Distributed Training, End to End
- Design a Feature Store
- A Write-Ahead Log You Can Implement in an Hour
- The GIL, Honestly
- Design a Stream Processor with Exactly-Once
- Design a Distributed Cache
- Back-of-the-Envelope: the Numbers That Design Systems
- Threads, Locks & the Anatomy of a Race
- Two Writers, One Row
- Consistency, Quorums & CAP
- How to Design a System in 60 Minutes
- Preparing for the Agent-Infrastructure Interview
- Rate Limiting: Four Algorithms, Honestly Compared
- Idempotency & the Exactly-Once Illusion
- Design Ad-Click Aggregation (Lambda vs Kappa)
- Design a DAG Job Scheduler
- Unique IDs at Scale
- From One Server to Millions of Users
- Consistent Hashing
- The Shuffle: How a Cluster Moves a Join
- Design a Multi-Tenant Query Engine on Object Storage
- Design a Durable Key-Value Store
- Design a Metrics & Monitoring System
- Design S3-Like Object Storage
- Design Search Autocomplete
- Design a Web Crawler
- Design a Distributed Message Queue
- Design a News Feed
- Design a File-Sync Service
- Design Google Maps
- Loading a Safetensors Checkpoint on a Multihost Cluster, Line by Line
- Model Surgery: Rewriting a Checkpoint’s Parameters, Line by Line
- CPython’s queue.py, Line by Line
- How Redux createStore Works, Line by Line
- functools.lru_cache, Line by Line
- Python’s Dict, Under the Hood
- Python Gotchas That Fail Interviews
- Iterators, Generators & the Dunder Protocols
- Errors, Edge Cases & Testing in the Room
- Designing a Clean Python API Under Pressure
- Idiomatic Python
- Writing Python Interviewers Trust
- The RL Cluster: Five Roles, One Mesh Dial
- Anatomy of a Production JAX LLM Trainer
- jit(grad(vmap(f))): Why Transform Order Changes the Answer
- Params Are Data: the Functional Model Behind Flax
- Design a Tool-Calling Platform for AI Agents
- The Tool Router: the Right 5 Tools out of 10,000
- Design a Webhook & Trigger Delivery Platform
- The OOD Interview, Decoded
- Three OOD Classics, Worked
- The Behavioral Interview, Decoded
- Stories That Carry Signal
- The Story Bank
- Buy the Cheapest Book
- The Seam Interface: computation_client.h, Line by Line
- The Lazy Tensor: What Happens Between Your Op and sync()
- How PJRT_DEVICE Becomes a Client: pjrt_registry.cpp, Line by Line
- The PJRT Boundary: One Training Step, Crossing by Crossing
- Writing torch_xla Notebooks That Survive Colab
- torch.compile Meets the Lazy Tensor: dynamo_bridge.py, Line by Line
- Where torch_xla Calls PJRT: pjrt_computation_client.cpp, Line by Line
- SPMD in torch_xla: A Mesh, an Annotation, and One Virtual Device
- From Pending IR to a Device Buffer: xla_graph_executor.cpp, Line by Line
- The Loop That Runs Every XLA Pass: hlo_pass_pipeline.cc, Line by Line
- The PJRT Plugin Contract: pjrt_api.cc, Line by Line
- HLO Module Anatomy: What XLA Holds While It Compiles
- Reading an XLA Dump: What the Compiler Writes, and How to Read It
- How a StableHLO Module Becomes HLO: mlir_to_hlo.cc, Line by Line
- Layout Assignment: Where the Copy in Your HLO Dump Comes From
- Buffer Assignment: How Every Value Gets an Address
- What Shares a GPU Kernel: priority_fusion.cc, Line by Line
- Instruction Fusion Legality: instruction_fusion.cc, Line by Line
- IFRT Arrays: pjrt_array.cc, Line by Line
- The CPU Thunk Executor: thunk_executor.cc, Line by Line
- The GPU Codegen Path: From a Fused HLO to a Kernel
- XLA Collectives: One Guest List, Then One Order
- Writing Pallas Kernels in Colab: What Interpret Mode Proves, and What It Cannot
- pallas_call, Line by Line: One Call, Three Calling Conventions
- The Mosaic TPU Pipeline, Line by Line
- Pallas on a TPU: From a Jaxpr to the Mosaic Dialect
- BlockSpec and the Grid: How Pallas Cuts an Array Into Kernel-Sized Pieces
- Pallas on a GPU: Two Backends, and What Each One Lets You Write
- Flash Attention on a TPU, Line by Line
- Splash Attention: When the Mask Stops Being Arithmetic and Becomes the Loop
- The tokamax Op and Its Autotuner, Line by Line
- Attention for sm90, Line by Line
- What JAX Hashes Before It Decides Not to Compile: cache_key.py, Line by Line
- Partial Evaluation: Splitting One Traced Call Into Two Programs
- linear_util.py, Line by Line
- Writing JAX Notebooks That Show Their Work
- Dispatch and Compilation: The Six Caches Under a jit Call
- Running a Tunix Recipe: From a YAML File to an RLCluster
- rl_cluster.py, Line by Line
- The RL Learner Loop, Line by Line
- The GRPO Learner, Line by Line
- How JAX Builds a Backward Pass Out of a Forward One: ad.py, Line by Line
- Where vmap Puts the Axis: batching.py, Line by Line
- How Every JAX Transform Unpacks Your Data: tree_util.py, Line by Line
- Control Flow as Primitives: One Trace, However Many Iterations
- The Tunix Sampler, Line by Line: Generation as Two Compiled Functions
- Rollout Backends and Weight Sync: Getting New Weights Into a Sampler
- The Trajectory Collect Engine, Line by Line
- The Tunix PEFT Trainer, Line by Line
- shard_map Internals: The Per-Device Body and the Checks Around It
- From Jaxpr to StableHLO: One Rule Per Primitive