Explore an AI system by changing its variables.
18 calculators, visualizers, evaluators and explorers for answering technical questions with explicit assumptions. Each tool documents its method, preserves source provenance and makes the scenario reproducible.
Available tools
A tool appears here when the English and Spanish versions share the same logic, sources and tests.
LLM cost and latency
Turn tokens, caching, traffic, TTFT, generation speed and concurrency into cost per request, monthly spend, response time and approximate capacity.
Open calculator →Model price and performance
Compare workload cost with Intelligence Index, output speed, TTFT and context while keeping pricing and performance provenance separate.
Open explorer →Inference VRAM
Separate weights, KV cache and runtime reserve to estimate total memory, memory per GPU, maximum context and approximate concurrency.
Open calculator →KV cache and context window
Visualize how KV-cache memory and context capacity change with GQA/MQA, precision, concurrency and a configurable memory budget.
Open explorer →Transformer attention
Manipulate scores, causal masking, softmax and values to see how one attention head turns a query into a weighted mixture.
Open visualizer →Token and context budget
Allocate the window across instructions, tools, history, RAG, the current message, output and safety headroom to surface overflow and future pressure.
Open planner →RAG retrieval
Measure Precision@k, Recall@k, MRR and nDCG on a visible ranking; test reranking while keeping retrieval quality separate from chunking and overlap footprint.
Open lab →RAG evaluation
Separate context relevance, faithfulness, correctness and coverage; inspect claims, uncertainty intervals and weights without hiding diagnosis inside one score.
Open evaluator →Voice-agent latency
Break down time to first audio across transport, turn end, STT, model, TTS and buffering, then calculate the interruption path separately.
Open explorer →Voice-agent cost and capacity
Turn calls, minutes, STT, TTS and model tokens into monthly spend, then size workers and provider limits without conflating different concurrency limits.
Open planner →Agent reliability and evaluation
Separate final success, first-pass success, retry recovery, tool decisions, timeouts and trajectory efficiency; apply explicit release gates.
Open evaluator →Prompt-injection threat paths
Trace paths from untrusted content into sensitive data, write-capable tools, external egress and persistent memory, then test which architectural boundaries cut each path.
Open explorer →Benchmark reliability
Check statistical resolution, saturation, sensitivity to invalid or potentially exposed items, and whether the ranking changes when task composition shifts.
Open explorer →Model capability timeline
Follow published results for one benchmark at a time, keep conditions and point-level provenance visible, and surface protocol breaks.
Open explorer →Scaling laws
Reallocate a fixed training budget between parameters and tokens, compare the optimum of a Chinchilla-style surface, and test sensitivity to fitted exponents.
Open explorer →Training compute and energy
Turn accelerators, MFU, schedule, average power and PUE into model FLOPs, estimated runtime, facility power and energy.
Open calculator →Datacenter AI capacity
Compare total power, PUE, physical slots, per-rack electrical power and cooling to find how many accelerators can be active and which physical constraint binds.
Open explorer →Global AI ecosystem
Compare investment, company formation, infrastructure, model development, talent and policy capacity with visible coverage, normalization and weights.
Open explorer →