Explore an AI system by changing its variables.
Calculators, visualizers and labs for answering AI engineering questions with explicit assumptions. Each tool documents its method, cites its sources and makes the scenario reproducible.
Available tools
A tool appears here when the English and Spanish versions share the same logic, sources and tests.
LLM cost and latency
Turn tokens, caching, traffic, TTFT, generation speed and concurrency into cost per request, monthly spend, response time and approximate capacity.
Open calculator →Model price and performance
Compare your workload cost with Intelligence Index, output speed, TTFT and context while keeping pricing and performance provenance separate.
Open explorer →Inference VRAM
Separate weights, KV cache and runtime reserve to estimate total memory, memory per GPU, maximum context and approximate concurrency.
Open calculator →KV cache and context window
Visualize how KV-cache memory and context capacity change with GQA/MQA, precision, concurrency and a configurable memory budget.
Open explorer →Transformer attention
Manipulate scores, causal masking, softmax and values to see how one attention head turns a query into a weighted mixture.
Open visualizer →Token and context budget
Allocate the window across instructions, tools, history, RAG, the current message, output and safety headroom to surface overflow and future pressure.
Open planner →RAG retrieval
Measure Precision@k, Recall@k, MRR and nDCG on a visible ranking; test reranking while keeping retrieval quality separate from chunking and overlap footprint.
Open lab →What comes next
The section grows one tool at a time. There are no empty placeholder pages: an entry appears when it has real interaction, ES/EN parity, provenance and tests.