Groq

Very fast inference for open models on custom LPU hardware.

Ask your AI about this, with this page as the source:ChatGPT ↗Claude ↗Perplexity ↗

Groq is an inference cloud. It runs open-weight models made by others, such as OpenAI's GPT-OSS, Qwen, Whisper and Orpheus text-to-speech, on its own LPU chips instead of GPUs, and sells access through the GroqCloud API. The point is low latency and high output speed, which matter for voice agents, live autocomplete and multi-step agent loops where every call adds wait time.

The developer model is an OpenAI-style HTTP API at api.groq.com/openai/v1. You can keep the OpenAI client libraries and change only the key and base URL, or use Groq's own Python and TypeScript SDKs. It supports Chat Completions and a Responses API, tool use with built-in browser search and code execution, remote MCP tools, structured outputs and prompt caching.

For a small team the relevant pieces are a free plan to start, a Batch API and Flex processing on the paid Developer plan, spend limits, and data controls an admin can set to zero retention. Groq runs it as a hosted service from data centers in North America, Europe, the Middle East and Australia; stored customer data sits in US Google Cloud buckets.

The limit is the catalog. You get only the models Groq has chosen to host, a few of them, such as the Llama 3 models, now listed as enterprise-only. Preview models can be withdrawn at short notice, and free-plan rate limits are low.

Where it fits

How Groq itself is built

4 tools, from its own code, website and Product Hunt page.

Who uses it

39 makers' products, each linked to the source that shows it, and 62 open-source projects that declare it in their code.

The maker says so 27Subprocessor list 13Declared in code 63How evidence is collected →
VectorizeAgents that remember. Agents that learn.

“We love using Groq hosting and models. They are a core part of our chat and widget agents. The models are lightning fast and the Groq team is fantastic.”

LLM APIMaker says so +1 · source ↗
MindPalMindPal is a no-code platform for turning a person's expertise into AI agents and multi-…

“You can select models hosted on Groq to power the AI agents you build on MindPal!”

LLM APIMaker says so · source ↗
graph8Turn buyer signals into outbound that works

“We use Groq for ultra-fast inference when analyzing millions of contact records and enriching them with AI. It enables us to run deep research and structured reasoning at speeds that would be impossible on standard GPU setups. This level of performance lets us deliver intelligent outputs in real time, even at scale. We're grateful to the Groq team for building the kind of infrastructure that makes this possible.”

LLM APIMaker says so · source ↗
AntispaceAction-Oriented AI: Translate Human Thoughts into Actions

“Groq Chat enhances our LPU's performance significantly, enabling faster inference and improved user interaction.”

LLM APIMaker says so · source ↗
KushoAIKushoAI uses AI agents to generate and run tests for web interfaces and backend APIs. It…

“Blazingly fast Large Language Model for code generation”

LLM APIMaker says so · source ↗
Yap-ItTurn messy voice notes into ready-to-publish content.

“YapIt only exists because of Groq. We needed instant transcription and rapid LLM formatting to make the app feel magical. Groq's whisper model and AI inference speeds are unmatched—they make our voice-to-text pipeline feel like it has zero latency.”

LLM APIMaker says so · source ↗
MagineSpawn vision-enabled AI agents autonomously browsing the web

“Magine runs on Groq's LLM inference for parallel context optimization.”

LLM APIMaker says so · source ↗
NBotPersonalized curators that surface what you care about

“Built with Groq for blazing-fast inference NBot is powered by Groq’s inference infrastructure, allowing us to process large volumes of content, run real-time summarization, and deliver high-quality AI responses at low latency. As we scale AI-powered feeds, chat, and synthesis across the internet, Groq’s performance and reliability have been critical to making the experience feel fast, responsive, and production-ready. Huge thanks to the Groq team for supporting us as an early partner.”

LLM APIMaker says so · source ↗
Voicr for MacDictate and get improved or translated text

“Shoutout to Groq — the lightning fast inference engine running Voicr's AI under the hood. Without Groq's speed, that under 3 seconds promise wouldn't be possible.”

LLM APIMaker says so · source ↗

Open source: a project that declares Groq as a dependency in its public code — verifiable, but not necessarily a live product.

What makers say

17 makers on why they use Groq, in their own words on Product Hunt.

Built with Groq for blazing-fast inference NBot is powered by Groq’s inference infrastructure, allowing us to process large volumes of content, run real-time summarization, and deliver high-quality AI responses at low latency. As we scale AI-powered feeds, chat, and synthesis across the internet, Groq’s performance and reliability have been critical to making the experience feel fast, responsive, and production-ready. Huge thanks to the Groq team for supporting us as an early partner.
NBotSep 2026 ↗
We use Groq for ultra-fast inference when analyzing millions of contact records and enriching them with AI. It enables us to run deep research and structured reasoning at speeds that would be impossible on standard GPU setups. This level of performance lets us deliver intelligent outputs in real time, even at scale. We're grateful to the Groq team for building the kind of infrastructure that makes this possible.
graph8Sep 2026 ↗
YapIt only exists because of Groq. We needed instant transcription and rapid LLM formatting to make the app feel magical. Groq's whisper model and AI inference speeds are unmatched—they make our voice-to-text pipeline feel like it has zero latency.
Yap-ItSep 2026 ↗
Shoutout to Groq — the lightning fast inference engine running Voicr's AI under the hood. Without Groq's speed, that under 3 seconds promise wouldn't be possible.
Voicr for MacSep 2026 ↗
We love using Groq hosting and models. They are a core part of our chat and widget agents. The models are lightning fast and the Groq team is fantastic.
VectorizeSep 2026 ↗

Loved and watch-outs

Themes that recur in makers' words and Hacker News comments, each linked to what it summarises, with how Product Hunt tags its reviews.

Most loved
  • Inference is extremely fast, enabling real-time voice and agent workflows. PHPH 2PH 3PH 4
  • The free tier is generous, with many models to choose from. PHPH 2HN
  • Its hosted Whisper gives near-instant transcription. PHPH 2HNPH 3
  • The team is helpful to work with. PHPH 2

Who switches

Public pull requests on GitHub since Oct 2024 whose title says "X to Y" — real code changes moving a project from one tool to another, by developers in general. Open a row to see the pull requests.

Gemini API → Groq152 PRs
OpenAI API → Groq50 PRs
Claude → Groq15 PRs
DeepSeek → Groq9 PRs
Mistral AI → Groq5 PRs
Groq → Gemini API64 PRs
Groq → OpenAI API25 PRs
Groq → Claude10 PRs
Groq → DeepSeek7 PRs
Groq → Mistral AI3 PRs

Questions makers ask about Groq

Can I use the OpenAI SDK with Groq?

Yes. Set the base URL to https://api.groq.com/openai/v1 and use your Groq API key. A few fields, such as logprobs, logit_bias and messages[].name, return a 400 error, and n must be 1. source ↗

Does Groq store my prompts and outputs?

Not by default for inference. It may log them for up to 30 days to troubleshoot reliability problems or investigate abuse, and batch and fine-tuning files are kept while those features need them. Any customer can turn on Zero Data Retention in Data Controls, which disables the features that need storage. source ↗

Where is my data stored?

Customer data Groq retains is kept in Google Cloud buckets in the United States. Transfers from other countries can rely on standard contractual clauses. source ↗

How strict are the free plan's limits?

Limits are per organization and measured in requests and tokens per minute and per day. On the free plan they are low, for example 30 requests per minute and 1,000 per day for the GPT-OSS models. The Developer plan raises them and adds Batch and Flex processing. source ↗

Is there a batch mode for bulk jobs?

Yes. Batch jobs cost 50% less than synchronous calls, don't count against your normal rate limits, and run within a window of 24 hours to 7 days. source ↗

Can I rely on any model in the catalog for production?

Only on the ones Groq labels as production models. Preview models are meant for evaluation and can be discontinued at short notice. source ↗

Which languages have an official SDK?

Python and JavaScript/TypeScript. For anything else, use the REST API or an OpenAI-compatible client. source ↗

Is Groq free?

Yes — there is a free tier a small product can run on; paid use starts at Pay per token. source ↗

Is Groq open source or self-hostable?

Not open source, and hosted only.

Can AI coding agents work with Groq?

No llms.txt, official MCP server or CLI found yet.

Who uses Groq?

39 makers' products we track, each with a source, and 62 open-source projects declare it in their code. source ↗