Fast hosted AI inference with model-specific limits on a free API tier.
Verified free allowance
Free-tier limits vary by model; current GPT-OSS 20B and 120B limits include 30 requests per minute, 1,000 requests per day, 8,000 tokens per minute, and 200,000 tokens per day.

GroqCloud provides hosted inference for supported language, audio, and compound AI models through an API. The free tier is useful for prototypes and learning, but limits apply at the organization level and a request can be blocked when any request, token, or audio allowance is reached.
A fast way to prototype against hosted open models for free, as long as the application can handle model-specific rate limits.
Paid plans: The Developer tier uses paid usage and higher limits; adding a payment method is required to upgrade.
These links support the important free-offer claims on this page.
Groq publishes model-specific Free Plan request, token, and audio limits and states that limits apply at the organization level.
Groq documents a Free tier and requires a payment method only when upgrading to the paid Developer tier.