Current Section

Overview

0%

← Back to AI Engineering Foundations
AI Engineering Foundations · Chapter 2

AI APIs & SDKs in Production

Turn foundation models into production software. Master HTTP token streaming, rate-limit resilience, multi-provider gateways, and bulletproof credential handling for scalable AI backends.

SSE StreamingToken Rate LimitingModel GatewaysCost Attribution

The Systems Law

An AI API is not a standard REST endpoint. It is a stateful, non-deterministic streaming stream over HTTP.

Traditional APIs return 200 OK responses in 50 milliseconds. LLM APIs take 500ms just to generate the first token and can run for 30 seconds. Production AI engineering requires handling long-lived HTTP connections, Server-Sent Events, token budgeting, and automated provider fallbacks when downstream inference clusters face capacity limits.

App Request → Gateway Auth & Quota → Streaming Transport → Token Chunking → Parsed State

Transport Architecture

The Client Request Lifecycle

Every production AI invocation moves through five explicit transport phases:

01

Client Initialization

Instantiate authenticated SDK client with scoped credentials, timeout limits, and connection pools.

02

Payload Assembly

Serialize role messages, tool definitions, temperature parameters, and response schema constraints.

03

Transport & Streaming

Transmit HTTPS payload over TLS and process chunked token deltas via Server-Sent Events (SSE).

04

Circuit Breaker & Retry

Intercept HTTP 429/503 errors and execute randomized jitter backoff before cascading failures.

05

Telemetry & Cost Audit

Log prompt tokens, completion tokens, latency to first token (TTFT), and tenant-level cost allocations.

Low-Latency UX

Token Streaming with Server-Sent Events (SSE)

Waiting for a 1,000-token completion to finish causes 8–15 seconds of blank UI latency. Production backends stream token chunks directly to users over HTTP/2 using Server-Sent Events:

python_streaming_fastapi.py
from fastapi import FastAPI
from fastapi.responses import StreamingResponse
from openai import AsyncOpenAI

app = FastAPI()
client = AsyncOpenAI()

async def token_generator(prompt: str):
    response = await client.chat.completions.create(
        model="gpt-4o",
        messages=[{"role": "user", "content": prompt}],
        stream=True,
    )
    async for chunk in response:
        delta = chunk.choices[0].delta.content or ""
        if delta:
            yield f"data: {delta}\n\n"

@app.post("/api/stream")
async def stream_ai(payload: dict):
    return StreamingResponse(
        token_generator(payload["prompt"]),
        media_type="text/event-stream"
    )
Key Metric: Time to First Token (TTFT). Streaming reduces perceived latency from 10,000ms down to ~350ms, keeping web and mobile users engaged.

Reliability Engineering

Handling HTTP 429s & Rate Limits

AI providers throttle based on two metrics: Requests Per Minute (RPM) and Tokens Per Minute (TPM). Unhandled rate spikes result in dropped calls. Production systems use exponential backoff with full randomized jitter:

Naive Immediate Retry (Thundering Herd)

Call 1 → 429 Rate Limit
Call 2 (Immediate) → 429
Call 3 (Immediate) → 429
Outcome: IP banned or connection throttled for 10 minutes.

Exponential Backoff with Jitter

Wait = Base * (2 ^ attempt) + random(0, 1)
Attempt 1: Wait ~2.1s
Attempt 2: Wait ~4.3s
Attempt 3: Wait ~8.7s
Outcome: Call resolves once token bucket replenishes.

Infrastructure Options

Architectural Patterns: SDKs vs Model Gateways

Choose the right connection topology based on your team's scale and compliance requirements:

Native FeaturesDirect Provider SDKs+

Using official libraries (`openai`, `@anthropic-ai/sdk`, `@google/genai`) gives immediate day-one access to provider-specific features like prompt caching, computer use, and native audio modalities.

Architectural Tradeoff

Tightly couples your application code to one vendor API schema, making multi-model routing harder.

Multi-Model OrchestrationUnified AI Gateways (LiteLLM, Portkey)+

A reverse proxy that maps the OpenAI request schema to 100+ backend LLMs. Enables automatic fallbacks, load balancing across API keys, and centralized token spend auditing.

Architectural Tradeoff

Adds an extra microservice hop and may lag behind bleeding-edge proprietary model features.

Enterprise ComplianceCloud Managed Endpoints (AWS Bedrock, Azure AI)+

Consume models through standard cloud IAM roles without external API keys. Keeps network traffic inside your private VPC, meeting strict HIPAA, SOC2, and data residency requirements.

Architectural Tradeoff

More complex IAM permission policies and occasional quota throttles compared to direct endpoints.

Security Boundary

Credential Storage & Zero Client Exposure

AI provider API keys carry direct financial and data liabilities. Leaked keys get scraped from GitHub within 60 seconds and abused for unauthorized training:

Never Call from Browsers

Never pass `OPENAI_API_KEY` into React, Next.js client components, or mobile binaries. Any user opening DevTools can extract your secret.

Use Ephemeral Backend Proxies

Route all browser requests through your internal Next.js `/api/chat` route handler. Authenticate user sessions before dispatching model queries.

Avoid

Common Production API Mistakes

Infinite Client Timeout

Leaving HTTP timeouts default causes thread pools to exhaust when upstream provider clusters experience degraded response times.

Blind Multi-Turn History

Passing entire conversation histories without token-pruning quickly pushes requests past model context limits and inflates operational costs.

Single API Key Bottlenecks

Funneling production, staging, and developer testing through one API key leads to unexpected tier rate-limit throttling in production.

Unbounded Payload Sizes

Allowing users to upload 500-page PDFs without chunking triggers HTTP 413 (Payload Too Large) and crashes backend memory.

Missing Cost Attribution

Failing to log token usage by tenant ID makes it impossible to detect abusive customers or calculate per-user gross margins.

Ignoring Context Caching

Re-sending large static instructions without exploiting prompt caching burns 10x more budget on every single request.

Release Gate

Pre-Production AI API Checklist

✓API keys are stored exclusively in environment variables or KMS, never bundled in client code.
✓All user-facing endpoints implement Server-Sent Events (SSE) streaming to minimize Time to First Token (TTFT).
✓HTTP 429 (Rate Limit) and HTTP 5xx errors are caught with exponential backoff and randomized jitter.
✓Per-request timeouts are configured (defaulting to 30s max for synchronous completions).
✓Token consumption (prompt, completion, cached tokens) is tracked and logged per tenant for cost attribution.
✓Model gateways or fallback providers are configured to prevent downtime during vendor API outages.
✓Backend microservices strip internal infrastructure details before returning error payloads to the client.

Key Takeaways

APIs turn models into resilient backend infrastructure.

Building with AI APIs requires moving past simple completion scripts. Stream token streams for responsive user interfaces, defend endpoints with exponential backoff, isolate credentials behind backend microservices, and instrument every call with cost and latency telemetry.

Secure Key Storage → SSE Stream Delivery → Jitter Backoff → Telemetry & Cost Accounting.