Skip to content

Saylek Reports

Keeping agent conversations cached

On NInfer, a long coding-agent conversation could lose its cache once the system-memory tier filled, so its next turn re-read 200K tokens from scratch. We fixed how the engine decides where that state fits.

Published
Topics
Speed
Setup
RTX 5090 · Qwen3.8-27B NVFP4 · NInfer
  • 113×

    faster turn after overflow

  • 99.996%

    of the prompt served from cache

  • Same

    speed when nothing overflows

The result

Resuming a 200K-token conversation after the system-memory tier fills. Seven 200K-token conversations fill the tier, two more arrive, and the first of those two is sent again. Stock: mean of two runs (47.4 s, 47.6 s). Fixed: the latest build; two earlier builds gave 0.40 s and 0.42 s.

Why a lost cache hurts

A cached turn only processes the new message. A lost cache means recomputing the entire prompt before the first token appears, and that cost grows with the conversation. Agent sessions reach 100K to 200K tokens within an hour of work.

Time to answer, by prompt length. Medians of four first sends (re-read) and four repeat sends (cached) per length, 16-token replies, on one RTX 5090.

What went wrong

NInfer keeps an active conversation's state on the GPU. When another conversation needs the GPU, that state moves to a tier of pinned system memory (RAM), so the next turn can resume instead of recomputing.

The system-memory tier is one memory arena shared by three allocation sizes: state images of 147 MiB, KV pages of 2.0 MiB and draft-model KV pages of 129 KiB. Each allocation needs one contiguous free range. The engine judged room by counting free bytes. After enough turnover the arena was fragmented, and in one traced failure it had 35.0 MB free against a 33.8 MB move, all of it in gaps smaller than a single KV page.

The engine approved the move without freeing anything, the move found nowhere to go, and the engine then released the conversation it had set out to preserve. From that point the tier kept its old conversations and lost every new one.

The fix

  • Room means placement. The arena simulates placing the exact allocations a move needs, after the ranges that evicting a candidate would return.
  • Eviction stops when the state fits. The engine evicts older conversations until the incoming state can be placed, prefers an eviction that completes placement over one that only frees bytes, and evicts nothing when no choice would make it fit.
  • One rule for every system-memory write. State captures and pause snapshots use the same placement check, and page reservations follow the same grouping as moves.

The tier recycles instead of freezing

Seven 200K-token conversations were sent in turn for 100 rounds, then an eighth arrived. During the rounds both builds served every turn from cache in under 0.5 s. The difference appears once the tier is full.

Cache state after an eighth conversation arrives. Stock keeps the seven old conversations indefinitely and re-reads the newest one on its next turn (47.2 s). The fix keeps the newest conversation (its next turn: 0.25 s) and evicts two older ones to make room for it.

Why 3 and 6 rather than 1 and 2: the fix evicts the conversations whose memory frees one contiguous range for the new one, so memory layout decides, not age. A block-based layout for this tier (#379) would remove the fragmentation and let the engine choose by recency and value instead, keeping the conversations an agent is most likely to return to.

In two further runs the fix held up across sizes and concurrency. With 16 conversations cycling through 16.5K, 66K, 131K and 200K tokens, every repeat send and 28 of 29 checks on an earlier conversation were served from cache. With two agents working at the same time on a full tier, every repeat send and recheck of the pair was served from cache in 0.23 to 0.64 s.

When it matters

The fix changes nothing while every active conversation fits on the GPU. On an RTX 5090 with this configuration the GPU holds about 254K tokens of KV cache, roughly one long agent session. The system-memory tier, and with it this fix, comes into play when work outgrows that:

  • an agent that runs sub-agents, or two agents working at once
  • switching between several long sessions or projects
  • long-running sessions where hours of turnover fragment the tier

Recommendations

  • If you run long agent sessions on NInfer

    Use a build with the fix: the saylek branch of Saylek-ai/ninfer today, or upstream once it merges. Give the system-memory tier as much pinned memory as the machine can spare with --host-context-mib; it decides how many evicted conversations stay resumable.

  • If you maintain NInfer

    The change is proposed upstream as three pull requests on #378. The remaining cost, evicting two conversations to admit one, comes from fragmentation itself. #379 proposes a block-based arena that would remove it.

  • What we measure next

    A two-agent replay on a full tier, stock and fixed on the same upstream base, to put a number on the gain for concurrent sessions in addition to the single-overflow case shown here.

Method

ItemDetail
HardwareOne NVIDIA RTX 5090, 50 GiB pinned system-memory tier
ModelQwen3.8-27B NVFP4 (qwen3_8_27b_nvfp4.ninfer), 240K-token context, two lanes, MTP speculative decoding
Stock buildNInfer 68c54356
Fixed builds68c54356 + fix; 81c8ce09 + fix
Common to allAn unrelated tool-call parsing fix (#244)
PromptsSynthetic, seeded random-word conversations sized with the engine's tokenizer; 16-token replies
TimingClient wall clock per request on the same machine
ReproductionScript and engine flags attached to #378

All figures come from synthetic runs on a test instance. The runs isolate one failure: the system-memory tier full and a conversation that must move to it. How often that happens in practice depends on how many long sessions share one GPU.

Sources

All Saylek Reports