Cerebras CS-4: a wafer-size chip runs an open model at 4,400 tokens a second
AnalysisSpeed, when you are waiting on an AI agent to grind through a long chain of steps, is the gap between a tool you keep open and one you close. Cerebras is selling that gap. On August 18 it launched the CS-4, a rack built from three dinner-plate-sized chips called the WSE-3 Turbo, each packing four trillion transistors, and clocked it at more than 4,400 tokens per second per user on GPT-OSS-120B, an open model from OpenAI. That is roughly 30 times a GPU setup and about ten times the work per watt of its own last machine. The product is patience, priced by the token.