What we actually ship
Use-case and data audit. Model selection. Prompt system. Retrieval over your corpus if the product needs it. Tools and guardrails. Usage metering and a fallback when the model is down. Handoff docs for your engineers.
Where cost hides
Tokens are yours. A chatty agent on a large corpus can cost more per month than the build. We put a budget alarm in before launch so that is not a surprise.
Preguntas
Do you fine-tune by default?
No. Retrieval plus a good prompt beats a fine-tune for most product FAQs. We fine-tune when the evals say the base model cannot get there.