High-performance inference you can reproduce.

Benchmark, optimize, validate, and deploy open models with the exact runtime evidence attached.

Qualification target AMD INSTINCT™ MI350X
1free-tier demo model
MI350X qualification target
request → scheduler → graphs → kernels → KV
Connecting to model…public demo
KIRA / PUBLIC MODEL RESPONSEready
Ask a question to run a real public inference request.
0 / 400
GPT-OSS 20B · zero-overage guard
FREE-TIER ONLY · 10 REQUEST RESERVATIONS / UTC DAY · PAID OVERAGE DISABLED

Designed around the production inference stack

ROCm
SGLang
vLLM
OPENAI API
AMD INSTINCT
NVIDIA

Benchmark claims should survive reproduction.

Kira records the exact hardware, software, model revision, request shape, failures, and raw measurements before a result becomes a headline.

MI350X qualification matrixPHYSICAL RUN PENDING
Matched serving sweepSGLANG · VLLM · KIRA PROFILE
WORKLOADTOKENSC1C8C32C64C128
Qwen 3.8 27B4K→512
Qwen 3.8 27B16K→512
DeepSeek V4 Flash4K→512
Long-context stress16K→256
○ physical measurement pending3 repetitions / point · errors retained

AMD Instinct MI350X

8× GPU

qualification target

Matched request pressure

C1 → C128

planned concurrency sweep

Promoted result requirement

fresh repetitions per point

Read benchmark plan
0

Unverified performance claims published.

Every future speedup has to map back to a retained run.

From workload to production without losing the evidence.

Benchmark the workload

Pin the model, request shape, runtime revision, token accounting, concurrency, and failure behavior.

Optimize the hot path

Change scheduler, graph, attention, KV, or kernel paths only when the complete service benefits.

Validate the result

Repeat the run, retain negative cells, and attach hashes for manifests, traces, and logs.

Deploy the profile

Carry the same validated runtime contract into managed or private infrastructure.

Request contractRUN / 01
MODELrevision + precision
SHAPEinput + output
LOADconcurrency + count
CONTROLmatched baseline
Runtime profilePROFILE / 02
scheduler.fill ██████████ 91%
graph.coverage █████████░ 86%
kv.pressure ██████░░░░ 61%
attention.path ████████░░ 79%
response.path █████░░░░░ 48%
Evidence receiptVERIFY / 03
run_id kira-mi350x-c64-r03
requests 384 / 384
client_errors retained
ooms retained
manifest_sha256 required
trace_sha256 required
Deployment contractDEPLOY / 04
Kira Cloudmanaged runtime profile
OpenAI-compatible endpoint
same
contract →
Kira Privateyour GPU fleet
same validated settings

One runtime contract. Choose where it runs.

Kira Cloud

Managed model serving with the benchmark profile and request surface already wired.

Run the public demo →
Kira Private

Bring the validated runtime profile into infrastructure you control.

Contact →

Engineering notes, not benchmark theater.

The useful record is the bottleneck, the attempted fix, the matched workload, the negative result, and the artifact that proves what ran.

MI350X / QUALIFICATION

Building the matched SGLang and vLLM serving matrix.

Scheduler capacity, graph coverage, KV pressure, exact request accounting, and end-to-end service measurements in one retained campaign.

Open methodology →
SCHEMA / EVIDENCE

A benchmark result should be machine-checkable.

The run schema records the hardware, software, workload, measurements, errors, and artifact hashes behind a result.

Open run schema →
RUNTIME / CAPACITY

KV cache is a deployment constraint, not a footnote.

Context length and concurrency can determine the viable serving shape before kernel tuning matters.

Discuss runtime →

Questions worth answering before deployment.

What is Kira Runtime?

+
Kira is a production inference layer focused on reproducible performance: benchmark the workload, qualify an optimized runtime profile, and carry that profile into deployment.

Is the public playground running on Kira's MI350X runtime?

+
No. The public demo is intentionally labeled and currently served through a tightly capped Cloudflare Workers AI free-tier path using GPT-OSS 20B. It automatically stops before Kira is allowed to consume paid inference, and its latency is never used as Kira benchmark evidence.

Why are the MI350X cells still empty?

+
Because the physical qualification run has not happened yet. Kira will not turn planned measurements into marketing numbers before the hardware run is complete and reproducible.

What has to be retained for a benchmark?

+
Hardware topology, driver and ROCm/CUDA versions, container/runtime revision, model revision and precision, workload shape, raw request measurements, errors/OOMs, and artifact hashes.

Can Kira target private infrastructure?

+
That is the intended deployment model: qualify the workload and runtime profile, then use the same contract on managed or customer-controlled GPU infrastructure.

Which accelerators are in scope?

+
The primary qualification focus is AMD Instinct MI350X-class infrastructure, with NVIDIA accelerator support also in scope. Exact model/runtime support must be validated per release.

Bring a real workload.

Share the model, context, concurrency target, latency constraints, and hardware. Kira starts from the exact serving contract.

Request qualification