by Arthur Rasmusson

Pod Efficiency Analyzer

The KV cache is your model's working memory — and recomputing it is GPU time you pay for twice. See what a memory fabric that anticipates is worth: lower TCO, more tokens per GPU, 10M-token context served warm, and prices your OpenRouter competitors can't follow.

Turn 2 of a 10.5M-token session Llama-4-Scout-17B-16E · 8×H100 · FP8 KV — measured
Inferra OFF
returning session · “Given everything in this repository, where is the race condition?”
time to first token
SESSION KV CACHE
10.5M tokens — ~26× larger than GPU memory
INFERRA — MEMORY MOVEMENT SCHEDULER HBM ↔ DRAM ↔ NVMe · one worklist
GPU · HBMscarce · compute-adjacent
GPU · ATTENTION KERNEL
EPHEMERAL
HOST DRAMTB-class · warm working set
STAGED
NVMe FABRICevery KV block · RDMA · durable
PERSISTENT
   328× faster
warm-reuse TTFT @ 10.5M tokens — 1 h 42 m → 18.7 s
more inference requests on the same GPUs
0%
lower power & infrastructure cost
0M
token context on Llama-4-Scout — served warm

Figures from Lightbits Labs' published benchmarks (the study), FarmGPU's independent KV-cache study (read it), and the Inferra GA measurements — see Benchmarks

Calculator 01

Total Cost of Ownership

Describe your serving fleet. We pull live on-demand pricing from the Infracost Cloud Pricing API and show what the same served load costs with and without Inferra KV-cache offload.

Fetching live price…
Multi-turn chat and agent workloads typically reuse 60–90% of prefill.
Current compute / month
With Inferra / month (incl. license)
Monthly savings
TCO reduction
Without Inferra
With Inferra
Monthly cost for the same served token volume.

Calculator 02

Inference Profit Calculator

Price your tokens with confidence. See your production TCO per million tokens with and without Inferra, your margin at a chosen list price, and how low you could price on OpenRouter while staying profitable.

Set this to the OpenRouter going rate for the model you serve.
Your fleet's real batched decode rate per GPU today. The published Inferra speedup ratio is applied on top of this figure.
Share of wall-clock time your fleet is actually generating tokens.
Fleet size, pricing, cache reuse, and license term carry over from the TCO Calculator tab.
Tokens/month without Inferra
Tokens/month with Inferra
TCO per 1M tokens without
TCO per 1M tokens with Inferra
Margin per 1M tokens (with Inferra)
Gross margin / month (with Inferra)
Lowest viable list price (30% margin)
Margin verdict at your price
Why this wins on OpenRouter: marketplaces rank providers on price, latency, uptime, and context length. Inferra cuts your cost per token so you can undercut, keeps turn-2 TTFT interactive so your latency stats shine, and lets you list context tiers competitors can't serve.

Calculator 03

Extended Context Advantage

Long-context, multi-turn sessions are unserveable with cache regeneration — users wait minutes for the first token. Pick your TTFT budget and see the maximum interactive context you can offer.

Max interactive context without Inferra
10,000,000
Max context with Inferra (Llama-4-Scout)
Context advantage over cache-regen serving
The pitch to your customers: most serving stacks cap interactive multi-turn context around 32K–128K tokens. With Inferra you can market full-document and agentic sessions at 1M tokens and beyond — and on Llama-4-Scout, whose architecture supports a 10-million-token context window, KV-cache offload is what makes those sessions practical to serve: an entire large codebase, a year of logs, or a long-lived agent's accumulated memory — analyzed in one conversation, restored warm in seconds on any node. A listing no cache-regen competitor can match.
See the measured benchmark results →

Evidence

Benchmark Results

Every number the calculators use, in one place — the Inferra GA measurement ladder, the published LightInferra study, and FarmGPU's independent verification. New results land here as they're measured.

Turn-2 TTFT — returning session, warm cache LOG SCALE

Llama-4-Scout-17B-16E, 8×H100, FP8 KV — measured on Inferra, reproducible recipe

100KTOKENS
7.4 s
191 ms
39×
1MTOKENS
2 m 03 s
1.66 s
74×
10.5MTOKENS
1 h 42 m
18.7 s
328×

Also measured: DeepSeek-R1-70B @ 102K context 53.4 s → 275 ms (194×); 8×H200 10M-token warm TTFT 11.6 s (601×); GTC demo config (Mode-5 VP) 8.3 h → 26 s (1,154×) at 10M. All results bit-identical to full recompute; reuse survives engine and fabric restarts.

Published turn-2 TTFT measurements

4× NVIDIA L40S, ScaleFlux PCIe 5.0 NVMe, RDMA 800G — Lightbits Labs study · FarmGPU independent study

ModelContextTTFT without Inferra TTFT with InferraSpeedup

More results are landing here.

8×H200 ladders, vLLM KV-Connector runs, and per-model token-economics sweeps are queued for this page. Want your configuration measured? Ask for a PoC →

Field tooling

Local GPU Analysis

For a precise, real-world snapshot, run our analyzer on one of your GPU servers. It inspects the system it runs on, estimates your Inferra uplift, and (with your consent) uploads the statistics to our API to enrich your PDF report.

$ curl -fsSL https://api.pod-efficiency.tools/sbin/gpu-analyzer.py | python3 -

Prefer to read it first? Download from api.pod-efficiency.tools/sbin/gpu-analyzer.py, then run python3 gpu-analyzer.py. Add --no-upload to keep results local only.

What it reads

Hostname, OS, CPU, RAM, NVIDIA GPUs (via nvidia-smi), NFS mounts, and disk capacity. Nothing else — the script is ~300 lines of stdlib Python you can audit.

What it tells you

A per-system report: detected GPUs, estimated tokens/s with and without Inferra, and how many fewer servers your workload would need.

Where data goes

One JSON document to api.pod-efficiency.tools/api/v1/stats over HTTPS. No credentials, no environment variables, no file contents.

Deliverable

Generate Your PDF Report

One click compiles a professional, boardroom-ready PDF from the numbers you've dialed in on the other tabs: executive summary, serving capacity, TCO, extended-context advantage, and token pricing guidance.

The server re-computes every figure with live Infracost pricing at generation time, so the PDF cites its own price source.

Next step

Talk to Lightbits Labs

Ready to turn these estimates into a proof of concept and a formal quote? Reach the Lightbits Labs team directly — the button below pre-fills an email with the fleet and workload you've dialed in on the calculator tabs.

Sales & general inquiries

Email: info@lightbitslabs.com

USA & Canada: 1-866-614-9802

International: +1-408-547-4391

Or use the form on lightbitslabs.com/contact-us.

Email sales with my numbers

Opens your mail client with a summary of your fleet, workload profile, license term, and estimated savings — nothing is sent until you hit send.

Offices

USA: 1830 The Alameda, San Jose, CA 95126

Israel: 17 Atir Yeda St., Kfar Saba 4464313

What to ask for: an Inferra proof of concept on your workload, formal per-GPU license pricing for your commitment term, and sizing guidance for the NVMe cache layer behind your fleet.

Serve 3× more inference on the GPUs you already have.

Run the numbers, generate the boardroom PDF, then let's prove it on your workload.

Generate my PDF report Contact sales