TileRT sets a new benchmark on AgentX, delivering 469 tok/s with GLM-5.3 on AMD Instinct MI355X GPUs, surpassing NVIDIA by 100 tok/s.
TileRT has reached a new milestone in AI inference performance, achieving 469 tokens per second (tok/s) for single-user generation throughput on the AgentX benchmark. This result was delivered using GLM-5.3 on 8 AMD Instinct MI355X GPUs and outpaced NVIDIA’s GB300 NVL72 system by over 100 tok/s, according to data published on September 24, 2026.
The AgentX benchmark, developed by SemiAnalysis, focuses on real-world, long-horizon agent workloads that reflect the operational dynamics of coding agents. Unlike traditional benchmarks constrained by fixed prompt lengths, AgentX tests how systems handle multi-turn interactions, growing contexts, and dependencies between agent components over extended sessions.
TileRT’s performance stands out not only for its raw speed but also for its ability to maintain efficiency over long contexts. Testing showed that when input context expanded from 1,000 to 1 million tokens—a 1,000x increase—TileRT still retained about two-thirds of its short-context performance, delivering 425 tok/s at the maximum context length. This sustained throughput is critical in agentic AI applications, where tasks often involve dynamically growing contexts and interactive responses.
The TileRT team attributes these results to its optimization strategies tailored for AMD’s CDNA 4 architecture, which powers the MI355X GPUs. Key innovations include:
These innovations align particularly well with AMD’s hardware capabilities, such as larger register files, fine-grained memory control, and instruction-level optimization through inline assembly.
TileRT’s advancement underscores the growing importance of optimizing inference speed in AI deployments.
Source link







