Tools

Explore an AI system by changing its variables.

18 calculators, visualizers, evaluators and explorers for answering technical questions with explicit assumptions. Each tool documents its method, preserves source provenance and makes the scenario reproducible.

Available

Available tools

A tool appears here when the English and Spanish versions share the same logic, sources and tests.

Calculator · LLMs · 01

LLM cost and latency

Turn tokens, caching, traffic, TTFT, generation speed and concurrency into cost per request, monthly spend, response time and approximate capacity.

Open calculator →
Explorer · Models · 02

Model price and performance

Compare workload cost with Intelligence Index, output speed, TTFT and context while keeping pricing and performance provenance separate.

Open explorer →
Calculator · Infrastructure · 03

Inference VRAM

Separate weights, KV cache and runtime reserve to estimate total memory, memory per GPU, maximum context and approximate concurrency.

Open calculator →
Explorer · Infrastructure · 04

KV cache and context window

Visualize how KV-cache memory and context capacity change with GQA/MQA, precision, concurrency and a configurable memory budget.

Open explorer →
Visualizer · Architecture · 05

Transformer attention

Manipulate scores, causal masking, softmax and values to see how one attention head turns a query into a weighted mixture.

Open visualizer →
Planner · LLMs · 06

Token and context budget

Allocate the window across instructions, tools, history, RAG, the current message, output and safety headroom to surface overflow and future pressure.

Open planner →
Lab · RAG · 07

RAG retrieval

Measure Precision@k, Recall@k, MRR and nDCG on a visible ranking; test reranking while keeping retrieval quality separate from chunking and overlap footprint.

Open lab →
Evaluator · RAG · 08

RAG evaluation

Separate context relevance, faithfulness, correctness and coverage; inspect claims, uncertainty intervals and weights without hiding diagnosis inside one score.

Open evaluator →
Explorer · Voice · 09

Voice-agent latency

Break down time to first audio across transport, turn end, STT, model, TTS and buffering, then calculate the interruption path separately.

Open explorer →
Planner · Voice · 10

Voice-agent cost and capacity

Turn calls, minutes, STT, TTS and model tokens into monthly spend, then size workers and provider limits without conflating different concurrency limits.

Open planner →
Evaluator · Agents · 11

Agent reliability and evaluation

Separate final success, first-pass success, retry recovery, tool decisions, timeouts and trajectory efficiency; apply explicit release gates.

Open evaluator →
Explorer · Security · 12

Prompt-injection threat paths

Trace paths from untrusted content into sensitive data, write-capable tools, external egress and persistent memory, then test which architectural boundaries cut each path.

Open explorer →
Explorer · Evaluation · 13

Benchmark reliability

Check statistical resolution, saturation, sensitivity to invalid or potentially exposed items, and whether the ranking changes when task composition shifts.

Open explorer →
Data explorer · Models · 14

Model capability timeline

Follow published results for one benchmark at a time, keep conditions and point-level provenance visible, and surface protocol breaks.

Open explorer →
Explorer · Scaling · 15

Scaling laws

Reallocate a fixed training budget between parameters and tokens, compare the optimum of a Chinchilla-style surface, and test sensitivity to fitted exponents.

Open explorer →
Calculator · Infrastructure · 16

Training compute and energy

Turn accelerators, MFU, schedule, average power and PUE into model FLOPs, estimated runtime, facility power and energy.

Open calculator →
Explorer · Infrastructure · 17

Datacenter AI capacity

Compare total power, PUE, physical slots, per-rack electrical power and cooling to find how many accelerators can be active and which physical constraint binds.

Open explorer →
Data explorer · Ecosystem · 18

Global AI ecosystem

Compare investment, company formation, infrastructure, model development, talent and policy capacity with visible coverage, normalization and weights.

Open explorer →
Standard

What a 5sigmas tool must show

01
Visible assumptions
Editable estimates are kept separate from observed data. If a result depends on an assumption, you can change it.
02
Primary sources
Prices, limits and changing data carry provenance and a verification date. Stale values are not presented as current.
03
Reproducible output
When useful, a scenario can be shared or exported with its inputs, outputs and provenance intact.