What we measure
Task accuracy on a frozen set. p95 latency. Cost at expected volume. Tool-calling reliability if the feature is an agent. Refusal quality on questions it must not answer.
What we do not do
Hard-code a single vendor SDK through the app. The gateway is one module. Swapping GPT for Claude should not touch your billing page.
Preguntas
Can we bring our own keys?
Yes. Netlify AI Gateway, OpenAI, Anthropic, or Vertex. Your account, your invoice from the provider.