Groq
United StatesInference platform (custom silicon)
Custom LPU chips serving open models at extreme token speeds.
What they do
Groq (not to be confused with xAI's Grok) designs the Language Processing Unit, a deterministic chip architecture that serves models like Llama and Qwen at hundreds to thousands of tokens per second — often the fastest inference available anywhere.
