Anything you build here can use AI without you opening an account anywhere, pasting a key, or picking a provider. The Gateway is already wired into your project: ask for the capability, and the right model is chosen for it.
Breadth, not a token gesture
Not two models with a dropdown. The Gateway spans the frontier families — Claude, Gemini, GPT and Grok — alongside the strong open-weight Chinese models worth using: DeepSeek, GLM and Kimi. Reasoning models for hard thinking, fast models for edits, long-context models for whole codebases and documents, vision models for images and screenshots.
Beyond text
- Speech to text — high-accuracy transcription with timestamps, diarisation and translation, including Whisper-class and Gemini transcription.
- Text to speech — natural narration and character voices, from Gemini TTS to Fish Audio and other top-tier voice engines.
- Audio understanding — summarise a call, pull decisions out of a meeting, caption a video.
- Music and sound — background beds, stings and generated tracks for product video.
- Images — generation and editing on brand, plus reading text and layout out of screenshots.
- Embeddings and search — so your app can answer questions over its own content.
How you use it in your app
Describe the behaviour — "transcribe the voice note and summarise it", "read this receipt", "narrate the walkthrough in a warm voice" — and the agent wires the calls, handles streaming, retries and rate limits, and keeps the credentials server-side where they belong.
Cost you can see
Every AI call in your project is metered in credits and shown per message, so a chat feature's real cost is visible before your users find it. Set limits and the Gateway respects them.
Still wondering something?
The assistant answers from these pages only.