SEIS: Self-Evolving Inference Systems

University of Chicago

Let AI agents evolve the entire inference engine.

The idea

Autonomous optimization.
Across the whole inference stack.

Faster inference makes language models cheaper to serve. SEIS gives AI agents a working inference engine and lets them improve it end to end. Across successive sessions, agents investigate bottlenecks, rewrite code, test their changes, and build on the code and reports left by earlier sessions. Parallel chains can also exchange discoveries.

Starting from mini-sglang, the agents developed optimized GPU kernels, fused operators, improved runtime execution, and prompt-lookup speculative decoding. The agent stays fixed throughout; the engine's implementation evolves.

SEIS protocol: agents optimize and evaluate an engine, inherit code and reports, and work through independent branches, sequential chains, or synchronized chains.
Engine code and reports pass between sessions; synchronized chains share discoveries. Click to enlarge.

The result

An evolved engine.
A substantial speedup.

Qwen3-0.6B · NVIDIA H100Single-request inference

3.27×

throughput over the original engine

594 → 1,942 tokens / second

89 ms

mean request latency

down from 295 ms

3.08×

without speculation

from kernel and runtime changes

Natural-request throughput

tokens / second ↑
mini-sglang original
594
vLLM
633
SGLang
626
TensorRT-LLM
690
SEIS best engine
1,942
100 natural-length requests, with prefill included. Throughput is the median of three repetitions. Production configurations were selected on development data. See Figure 1 and Appendix E of the paper.

The best engine's observed accuracy changes were +1.06 percentage points on GSM8K and −0.37 on long-input retrieval, relative to the original engine.

What we learn

Better systems through accumulated work.

01 / Experience

Build on earlier sessions.

Inherited code and reports let agents combine improvements over time. In these experiments, sustained chains outperform independent attempts.

02 / Collaboration

Share useful discoveries.

Exchanging code across chains spreads working designs. The best engine combines its own kernels and runtime with a peer's decoding policy.

03 / Evaluation

Keep checking capability.

Speed and development checks tell only part of the story. Held-out task evaluation and code inspection expose failures that the optimization score misses.

Citation

BibTeX
@article{xu2026seis,
  title   = {{SEIS}: Self-Evolving Inference Systems},
  author  = {Xu, Zhen and Liu, Jingyu and Li, Zongze and
             Rabbani, Tahseen and Zhang, Ce},
  journal = {Preprint},
  year    = {2026}
}