High-performance inference you can reproduce.
Benchmark, optimize, validate, and deploy open models with the exact runtime evidence attached.
Designed around the production inference stack
Benchmark claims should survive reproduction.
Kira records the exact hardware, software, model revision, request shape, failures, and raw measurements before a result becomes a headline.
AMD Instinct MI350X
8× GPUqualification target
Matched request pressure
C1 → C128planned concurrency sweep
Promoted result requirement
3×fresh repetitions per point
Unverified performance claims published.
Every future speedup has to map back to a retained run.From workload to production without losing the evidence.
Pin the model, request shape, runtime revision, token accounting, concurrency, and failure behavior.
Change scheduler, graph, attention, KV, or kernel paths only when the complete service benefits.
Repeat the run, retain negative cells, and attach hashes for manifests, traces, and logs.
Carry the same validated runtime contract into managed or private infrastructure.
graph.coverage █████████░ 86%
kv.pressure ██████░░░░ 61%
attention.path ████████░░ 79%
response.path █████░░░░░ 48%
requests 384 / 384
client_errors retained
ooms retained
manifest_sha256 required
trace_sha256 required
OpenAI-compatible endpoint
contract →
same validated settings
One runtime contract. Choose where it runs.
Bring the validated runtime profile into infrastructure you control.
Contact →Engineering notes, not benchmark theater.
The useful record is the bottleneck, the attempted fix, the matched workload, the negative result, and the artifact that proves what ran.
Building the matched SGLang and vLLM serving matrix.
Scheduler capacity, graph coverage, KV pressure, exact request accounting, and end-to-end service measurements in one retained campaign.
Open methodology →A benchmark result should be machine-checkable.
The run schema records the hardware, software, workload, measurements, errors, and artifact hashes behind a result.
Open run schema →KV cache is a deployment constraint, not a footnote.
Context length and concurrency can determine the viable serving shape before kernel tuning matters.
Discuss runtime →Questions worth answering before deployment.
What is Kira Runtime?
+Is the public playground running on Kira's MI350X runtime?
+Why are the MI350X cells still empty?
+What has to be retained for a benchmark?
+Can Kira target private infrastructure?
+Which accelerators are in scope?
+Bring a real workload.
Share the model, context, concurrency target, latency constraints, and hardware. Kira starts from the exact serving contract.
Request qualification