← Back to community benchmarks

gpt-oss-120b-MXFP4-Q8

M5 Max (40c) · 128 GB · 4bit · 2026-04-04

Performance

32k

tokens

1,368

PP tok/s

39.0

TG tok/s

23958

TTFT (ms)

60.6

Peak mem (GB)

Hardware

Chip M5 Max (40c)

Memory 128 GB

GPU Cores 40

Software

oMLX v0.3.2

macOS macOS 26.4

Context 32,768