AI API Platforms

Platforms that give developers API access to AI models without running GPUs themselves. This note compares the categories and links to each vendor’s own pricing page. Prices are deliberately not copied here: they change monthly (for example, Anthropic’s Claude Sonnet 5 introductory price became the standard price in 2026), so always check the live page.

Categories

  • First-party model APIs: the lab serves its own models (openai, anthropic, Google).
  • Cloud marketplaces: hyperscaler serves third-party and own models with cloud billing and SLAs (Amazon Bedrock on amazon-web-services; Microsoft Foundry; Google’s Gemini Enterprise Agent Platform, the rebrand of Vertex AI announced 2026-04-22 at Cloud Next; see google-vertex-ai-agent-builder, microsoft-foundry-note).
  • Media/inference specialists: serverless endpoints for open and community models, billed per output or per second (fal, replicate, Kie.ai).
  • Gateways: sit in front of several providers (see PortKey).

Platforms

PlatformWhat it isPricing model (verified)Where to check prices
OpenAI APIFirst-party GPT modelsPer-token; Standard, Batch (about half of standard on most models), Flex (select models), Fast (the former Priority tier, renamed 2026-07-30) and Ultrafast (limited to gpt-6-astra)https://developers.openai.com/api/docs/pricing
Anthropic API (Claude)First-party Claude modelsPer-token; Batch API gives 50% off input and output; prompt-cache reads are 0.1x base input (0.025x on Fable 5.1, 0.05x on Opus 5.5); 1M-token context at standard pricing for Claude 4.6 and laterhttps://platform.claude.com/docs/en/about-claude/pricing
Gemini Enterprise Agent Platform (formerly Vertex AI)Google Cloud model and agent platform, 200+ models including ClaudePer model and servicehttps://cloud.google.com/vertex-ai/generative-ai/pricing
Amazon BedrockMulti-vendor models on AWSStandard (on-demand per token), Batch (50% lower for select models), Priority (+75%), Flex (-50%), Provisioned Throughput (hourly, 1- or 6-month commitment)https://aws.amazon.com/bedrock/pricing/
fal (fal-dot-ai)Generative media (video, image, audio, 3D)Per-output model APIs (video per second, image per image or megapixel, audio per characters or seconds, 3D per generation) or hourly GPU rentalhttps://fal.ai/pricing
replicateOpen and community models, custom models via Cog; part of Cloudflare since the acquisition closed 2025-12-01 (see replicate)Per-second hardware billing for most public models; some models priced per input/outputhttps://replicate.com/pricing
Kie.ai (kie-dot-ai)Aggregator for media and chat models, credit-basedPrepaid credits (secondary sources: credits do not expire; vendor page not retrievable, rate unverified)https://kie.ai/pricing

Anthropic models are also sold through Bedrock, Google Cloud, Claude Platform on AWS and Claude in Microsoft Foundry; the last two bill in “Claude Consumption Units” at the same per-model rates as the first-party API.

Choosing

  • Newest model features first, maximum control: first-party API.
  • Already committed to a cloud: Bedrock, Foundry or Google’s agent platform for IAM, networking, regional endpoints and procurement (Anthropic notes regional endpoints carry a 10% premium over global ones on Bedrock and Google Cloud).
  • Image/video/audio generation in products: fal or Replicate; compare output quality and per-output price per model, not the platform.
  • Custom model deployment: Replicate (Cog packaging) or fal custom deployments; self-hosting on Kubernetes is the alternative.
  • One key for many vendors, failover, cost tracking: a gateway such as PortKey (acquired by Palo Alto Networks, announced 2026-04-30, closed 2026-05-29 per press reports; part of its Prisma AIRS platform) or Cloudflare AI Gateway.

Cost levers (confirmed)

  • Batch APIs: 50% off at Anthropic and on Bedrock for select models; OpenAI shows a batch tier at about half of standard.
  • Prompt caching: Anthropic cache reads cost a fraction of base input price; discounts stack with batch.
  • Right-size hardware where the platform exposes GPU choice (Replicate lists T4 through H100 among others).
  • Committed capacity (Bedrock provisioned throughput, enterprise agreements) only for sustained load.

Not verified

Earlier drafts of this note claimed “60-70% cheaper”, “4x faster”, “5-10 s cold start” and “LoRA in under 5 minutes” for individual vendors. These are vendor marketing claims and were removed because they could not be confirmed.

replicate, fal-dot-ai, kie-dot-ai, openai, anthropic, google-vertex-ai-agent-builder, microsoft-foundry-note, PortKey, cloudflare

Sources