Post-Cutoff.com
  1. Home
  2. Posts
  3. RunInfra (YC F26)

RunInfra (YC F26)

RunInfra (YC F26) @runinfrai · x · 2026-09-30 · ★★★ · archived

Open the original ↗

Cited as a source by: glm-5-3-flash

Summary

Archived text

we spent september rewriting the kernels behind GLM 5.3 Flash. major release is live on RunInfra today

670 tok/s on Vercel AI Gateway

$0.11 per 1M input, $0.45 per 1M output, $0.03 per 1M cached. 1M token context. FP8, vendor-native release

it now runs on AMD. same model, same API, more capacity behind it

99.7% cache hit rate over the last 24 hours. every hit and miss shows up in your dashboard, so you can track the cache logs yourself

OpenAI-compatible chat completions and Anthropic-compatible /v1/messages. text and image input, tool calling, JSON mode, streaming

zero data retention. never used for training

https://runinfra.ai/inference-api/glm-5-3-flash

Media: https://pbs.twimg.com/media/HTcf8FcasAEYUEg.png?name=orig

views 204342 · likes 1382 · reposts 79 · replies 59 (at fetch time)

Archived 2026-09-30 via fxtwitter (unofficial).

Archived text

we spent september rewriting the kernels behind GLM 5.3 Flash. major release is live on RunInfra today

670 tok/s on Vercel AI Gateway

$0.11 per 1M input, $0.45 per 1M output, $0.03 per 1M cached. 1M token context. FP8, vendor-native release

it now runs on AMD. same model, same API, more capacity behind it

99.7% cache hit rate over the last 24 hours. every hit and miss shows up in your dashboard, so you can track the cache logs yourself

OpenAI-compatible chat completions and Anthropic-compatible /v1/messages. text and image input, tool calling, JSON mode, streaming

zero data retention. never used for training

https://runinfra.ai/inference-api/glm-5-3-flash

Media: https://pbs.twimg.com/media/HTcf8FcasAEYUEg.png?name=orig

views 204342 · likes 1382 · reposts 79 · replies 59 (at fetch time)

Archived 2026-09-30 via fxtwitter (unofficial).

All posts · id: x-runinfrai-2105186912634057086