AI API Platforms
Platforms that give developers API access to AI models without running GPUs themselves. This note compares the categories and links to each vendor’s own pricing page. Prices are deliberately not copied here: they change monthly (for example, Anthropic’s Claude Sonnet 5 introductory price became the standard price in 2026), so always check the live page.
Categories
- First-party model APIs: the lab serves its own models (openai, anthropic, Google).
- Cloud marketplaces: hyperscaler serves third-party and own models with cloud billing and SLAs (Amazon Bedrock on amazon-web-services; Microsoft Foundry; Google’s Gemini Enterprise Agent Platform, the rebrand of Vertex AI announced 2026-04-22 at Cloud Next; see google-vertex-ai-agent-builder, microsoft-foundry-note).
- Media/inference specialists: serverless endpoints for open and community models, billed per output or per second (fal, replicate, Kie.ai).
- Gateways: sit in front of several providers (see PortKey).
Platforms
| Platform | What it is | Pricing model (verified) | Where to check prices |
|---|---|---|---|
| OpenAI API | First-party GPT models | Per-token; Standard, Batch (about half of standard on most models), Flex (select models), Fast (the former Priority tier, renamed 2026-07-30) and Ultrafast (limited to gpt-6-astra) | https://developers.openai.com/api/docs/pricing |
| Anthropic API (Claude) | First-party Claude models | Per-token; Batch API gives 50% off input and output; prompt-cache reads are 0.1x base input (0.025x on Fable 5.1, 0.05x on Opus 5.5); 1M-token context at standard pricing for Claude 4.6 and later | https://platform.claude.com/docs/en/about-claude/pricing |
| Gemini Enterprise Agent Platform (formerly Vertex AI) | Google Cloud model and agent platform, 200+ models including Claude | Per model and service | https://cloud.google.com/vertex-ai/generative-ai/pricing |
| Amazon Bedrock | Multi-vendor models on AWS | Standard (on-demand per token), Batch (50% lower for select models), Priority (+75%), Flex (-50%), Provisioned Throughput (hourly, 1- or 6-month commitment) | https://aws.amazon.com/bedrock/pricing/ |
| fal (fal-dot-ai) | Generative media (video, image, audio, 3D) | Per-output model APIs (video per second, image per image or megapixel, audio per characters or seconds, 3D per generation) or hourly GPU rental | https://fal.ai/pricing |
| replicate | Open and community models, custom models via Cog; part of Cloudflare since the acquisition closed 2025-12-01 (see replicate) | Per-second hardware billing for most public models; some models priced per input/output | https://replicate.com/pricing |
| Kie.ai (kie-dot-ai) | Aggregator for media and chat models, credit-based | Prepaid credits (secondary sources: credits do not expire; vendor page not retrievable, rate unverified) | https://kie.ai/pricing |
Anthropic models are also sold through Bedrock, Google Cloud, Claude Platform on AWS and Claude in Microsoft Foundry; the last two bill in “Claude Consumption Units” at the same per-model rates as the first-party API.
Choosing
- Newest model features first, maximum control: first-party API.
- Already committed to a cloud: Bedrock, Foundry or Google’s agent platform for IAM, networking, regional endpoints and procurement (Anthropic notes regional endpoints carry a 10% premium over global ones on Bedrock and Google Cloud).
- Image/video/audio generation in products: fal or Replicate; compare output quality and per-output price per model, not the platform.
- Custom model deployment: Replicate (Cog packaging) or fal custom deployments; self-hosting on Kubernetes is the alternative.
- One key for many vendors, failover, cost tracking: a gateway such as PortKey (acquired by Palo Alto Networks, announced 2026-04-30, closed 2026-05-29 per press reports; part of its Prisma AIRS platform) or Cloudflare AI Gateway.
Cost levers (confirmed)
- Batch APIs: 50% off at Anthropic and on Bedrock for select models; OpenAI shows a batch tier at about half of standard.
- Prompt caching: Anthropic cache reads cost a fraction of base input price; discounts stack with batch.
- Right-size hardware where the platform exposes GPU choice (Replicate lists T4 through H100 among others).
- Committed capacity (Bedrock provisioned throughput, enterprise agreements) only for sustained load.
Not verified
Earlier drafts of this note claimed “60-70% cheaper”, “4x faster”, “5-10 s cold start” and “LoRA in under 5 minutes” for individual vendors. These are vendor marketing claims and were removed because they could not be confirmed.
Related
replicate, fal-dot-ai, kie-dot-ai, openai, anthropic, google-vertex-ai-agent-builder, microsoft-foundry-note, PortKey, cloudflare
Sources
- Anthropic pricing, https://platform.claude.com/docs/en/about-claude/pricing (accessed 2026-10-02)
- OpenAI API pricing, https://developers.openai.com/api/docs/pricing (accessed 2026-10-02)
- Amazon Bedrock pricing, https://aws.amazon.com/bedrock/pricing/ (accessed 2026-10-02)
- fal pricing, https://fal.ai/pricing (accessed 2026-10-02)
- Replicate pricing, https://replicate.com/pricing (accessed 2026-10-02)
- Google Cloud, Introducing Gemini Enterprise Agent Platform, https://cloud.google.com/blog/products/ai-machine-learning/introducing-gemini-enterprise-agent-platform (accessed 2026-10-02)
- Kie.ai, https://kie.ai/pricing (page content not retrievable; accessed 2026-10-02; secondary: betterai.dev/kie)
- Palo Alto Networks / Portkey (Yahoo Finance, GovConWire, Futurum via search, accessed 2026-10-02)
- Gemini Enterprise Agent Platform rebrand, https://www.hpcwire.com/aiwire/2026/04/23/google-unveils-gemini-enterprise-agent-platform/ (accessed 2026-10-02)