OpenFabricAgentic system

OPENFABRIC / AI NETWORK SYSTEMS

Intelligence should move without limits.

We link full-system simulation to network hardware, so the next AI service can be predicted before it is deployed.

01 Model02 Simulate03 Measure04 Decide

The queue can cross its limit before control completes one round trip.

DCQCN controls each flow after congestion appears. Synchronized MoE traffic can fill the switch buffer first, then pay in marks, rate cuts, loss and tail latency.

Reference study: Kimi K3, 256 trays, 2,048 GPUs. Official model geometry, SimLLM M5 traffic mapping, declared capacity plan.

SIMLLM STUDY / 2026-08-05

Decode pool

D
Pool192 trays24 × 64-GPU replicas
Traffic / replica step302.5 GB256-request decode step
Synchronized pool wave7.26 TBinter-node payload, MXFP8
MoE operations18492 layers × dispatch + combine

PD KV-transfer traffic is not included because milestone M6 is still open. The stage pool count is derived from the visible 30 req/s, 8K input, 1K output capacity plan.

MEASURED MECHANISM REPLAYMoE dispatch → compute → combine0.18× conceptual playback · timing not to scale
STAGE 01 / 06Dispatch

Tokens fan out to experts across multiple paths. The offered load stays distributed and the destination queue remains below ECN Kmin. This calm fan-out is conceptual and precedes the isolated incast replay.

NO CONGESTION
8 GPU / tray64 KiB / source · 1.31 μs injection

Mechanism view 32-way convergence into a 1 MiB buffer at 400G. Animation timing is intentionally slowed and not to scale.

00.68 μs buffer full4 μs RTT
302.1×measured DCQCN p99 anchor
PACKET REPLAY / 2026-08-06

The burst is short. The recovery tail is not.

Packet-level HTSIM replays use identical aligned-start flows, 400G links, a 1 MiB shared buffer, PFC off and seed 1. They are simulation evidence, not measurements from a production Kimi deployment.

01Flow completion CDFDecode + prefill combine replays
Two flow completion CDFs comparing null-network ideal and RoCEv2 DCQCN for decode and prefill combine replays
Decode contains early DCQCN winners before 17 of 32 flows enter the 50 ms RTO tail, so the null curve is not claimed as a per-flow lower bound. The all-flow makespan below is the reviewable comparison.
02Null network vs RoCEv2/DCQCNTime until all 32 flows complete
Log-scale dumbbell comparison of null-network ideal and RoCEv2 DCQCN all-flow completion time for decode and prefill combine replays
Decode completes in 0.044 ms under the null network and 51.010 ms under DCQCN; prefill completes in 1.344 ms versus 355.009 ms. The comparator’s 50 ms silent-RTO quantizes the tail.Read the registered study ↗

Do not wait for congestion to explain itself.

Open the fabric.Expand the future.

Predict the burst. Schedule the packet. Prove the outcome.

Keep the scheduler real. Simulate the machine beneath it.

SimLLM runs the serving framework’s own batching and cache-control logic, then replaces model execution with calibrated GPU service and packet-level fabric simulation. Every scheduler decision becomes compute and collective work; network completion advances the virtual clock and changes the next decision.

DeterministicLine rateLossless by designOne design target, verified together
Explore the open-source SimLLM architecture
  1. 01Input

    Request trace

    Arrivals + token lengths

  2. 02Runs real

    Real scheduler

    vLLM / SGLang batching + KV control

  3. 03Contract

    StepRecord

    ExecutionLowerer + compute provider

  4. 04Simulated

    ExecutionGraph

    GPU, HBM, DMA and resource queues

  5. 05Causal

    Collective trace

    TP + MoE dispatch/combine → GOAL

  6. 06htsim

    Packet fabric

    Clos + RNIC + congestion control

  7. 07Feedback

    StepResult

    FCT → TTFT / TPOT / SLO

Side inputs
Placement manifestrank → node → GPU → NIC
Fabric topologyClos, links, routes and queues
03 / 04

MANAGE THE WHOLE SYSTEM

Performance is a chain of causes.

Measure where control arrives too late.

Model the gaps that averages cannot see.

Most tools stop at GPU time or bytes divided by bandwidth. OpenFabric keeps the scheduler, queue, GPU, NIC and packet in one mental model.

  1. Now / Workload

    Requests and queueing

    Arrival processes, prefix reuse and real frontend scheduling drive the request trace.

    TTFT / TPOT
  2. Now / Network

    Collectives and packets

    TP, MoE and flow-level work run directly through analytical and packet backends.

    FCT ledger
  3. Next / Runtime

    GPU, HBM, DMA, WQE and NIC

    Explicit resource ownership reveals queueing that a single compute duration cannot represent.

    Calibrated
  4. Next / Hardware

    First-principles packet scheduling

    Co-design the protocol and NIC for deterministic, line-rate, lossless operation.

    Built + proved
  5. Future / Agent

    One-click deployment loop

    Pre-study, calibrate, deploy, observe and retune with an agent operating the full system model.

    Explore the agentic system
    Closed loop

CURRENT DIRECTION

Cobalt / Ink

Closest to the current OpenFabric identity. Select another direction to recolor the complete page.