Calculator 01
Total Cost of Ownership
Describe your serving fleet. We pull live on-demand pricing from the Infracost Cloud Pricing API and show what the same served load costs with and without Inferra KV-cache offload.
Calculator 02
Inference Profit Calculator
Price your tokens with confidence. See your production TCO per million tokens with and without Inferra, your margin at a chosen list price, and how low you could price on OpenRouter while staying profitable.
Calculator 03
Extended Context Advantage
Long-context, multi-turn sessions are unserveable with cache regeneration — users wait minutes for the first token. Pick your TTFT budget and see the maximum interactive context you can offer.
Evidence
Benchmark Results
Every number the calculators use, in one place — the Inferra GA measurement ladder, the published LightInferra study, and FarmGPU's independent verification. New results land here as they're measured.
Turn-2 TTFT — returning session, warm cache LOG SCALE
Llama-4-Scout-17B-16E, 8×H100, FP8 KV — measured on Inferra, reproducible recipe
Also measured: DeepSeek-R1-70B @ 102K context 53.4 s → 275 ms (194×); 8×H200 10M-token warm TTFT 11.6 s (601×); GTC demo config (Mode-5 VP) 8.3 h → 26 s (1,154×) at 10M. All results bit-identical to full recompute; reuse survives engine and fabric restarts.
Published turn-2 TTFT measurements
4× NVIDIA L40S, ScaleFlux PCIe 5.0 NVMe, RDMA 800G — Lightbits Labs study · FarmGPU independent study
| Model | Context | TTFT without Inferra | TTFT with Inferra | Speedup |
|---|
More results are landing here.
8×H200 ladders, vLLM KV-Connector runs, and per-model token-economics sweeps are queued for this page. Want your configuration measured? Ask for a PoC →
Field tooling
Local GPU Analysis
For a precise, real-world snapshot, run our analyzer on one of your GPU servers. It inspects the system it runs on, estimates your Inferra uplift, and (with your consent) uploads the statistics to our API to enrich your PDF report.
curl -fsSL https://api.pod-efficiency.tools/sbin/gpu-analyzer.py | python3 -
Prefer to read it first? Download from
api.pod-efficiency.tools/sbin/gpu-analyzer.py,
then run python3 gpu-analyzer.py. Add --no-upload to keep
results local only.
Hostname, OS, CPU, RAM, NVIDIA GPUs (via nvidia-smi), NFS mounts, and disk capacity. Nothing else — the script is ~300 lines of stdlib Python you can audit.
A per-system report: detected GPUs, estimated tokens/s with and without Inferra, and how many fewer servers your workload would need.
One JSON document to api.pod-efficiency.tools/api/v1/stats
over HTTPS. No credentials, no environment variables, no file contents.
Deliverable
Generate Your PDF Report
One click compiles a professional, boardroom-ready PDF from the numbers you've dialed in on the other tabs: executive summary, serving capacity, TCO, extended-context advantage, and token pricing guidance.
The server re-computes every figure with live Infracost pricing at generation time, so the PDF cites its own price source.
Next step
Talk to Lightbits Labs
Ready to turn these estimates into a proof of concept and a formal quote? Reach the Lightbits Labs team directly — the button below pre-fills an email with the fleet and workload you've dialed in on the calculator tabs.
Sales & general inquiries
Email: info@lightbitslabs.com
USA & Canada: 1-866-614-9802
International: +1-408-547-4391
Or use the form on lightbitslabs.com/contact-us.
Opens your mail client with a summary of your fleet, workload profile, license term, and estimated savings — nothing is sent until you hit send.
Offices
USA: 1830 The Alameda, San Jose, CA 95126
Israel: 17 Atir Yeda St., Kfar Saba 4464313