RunInfra (YC F26)
RunInfra (YC F26) @runinfrai · x · 2026-09-30 · ★★★ · archived
Cited as a source by: glm-5-3-flash
Summary
Archived text
we spent september rewriting the kernels behind GLM 5.3 Flash. major release is live on RunInfra today
670 tok/s on Vercel AI Gateway
$0.11 per 1M input, $0.45 per 1M output, $0.03 per 1M cached. 1M token context. FP8, vendor-native release
it now runs on AMD. same model, same API, more capacity behind it
99.7% cache hit rate over the last 24 hours. every hit and miss shows up in your dashboard, so you can track the cache logs yourself
OpenAI-compatible chat completions and Anthropic-compatible /v1/messages. text and image input, tool calling, JSON mode, streaming
zero data retention. never used for training
Media: https://pbs.twimg.com/media/HTcf8FcasAEYUEg.png?name=orig
views 204342 · likes 1382 · reposts 79 · replies 59 (at fetch time)
Archived 2026-09-30 via fxtwitter (unofficial).
Archived text
we spent september rewriting the kernels behind GLM 5.3 Flash. major release is live on RunInfra today
670 tok/s on Vercel AI Gateway
$0.11 per 1M input, $0.45 per 1M output, $0.03 per 1M cached. 1M token context. FP8, vendor-native release
it now runs on AMD. same model, same API, more capacity behind it
99.7% cache hit rate over the last 24 hours. every hit and miss shows up in your dashboard, so you can track the cache logs yourself
OpenAI-compatible chat completions and Anthropic-compatible /v1/messages. text and image input, tool calling, JSON mode, streaming
zero data retention. never used for training
Media: https://pbs.twimg.com/media/HTcf8FcasAEYUEg.png?name=orig
views 204342 · likes 1382 · reposts 79 · replies 59 (at fetch time)
Archived 2026-09-30 via fxtwitter (unofficial).
All posts · id: x-runinfrai-2105186912634057086