Generative AI product development at scale means solving five problems that do not exist at pilot size: unit cost that grows with usage, evaluation that has to run as release infrastructure, model versions that get deprecated underneath you, tenant data isolation, and latency budgets that cannot absorb a retry.
None of these are model problems. All of them are product engineering problems, and they arrive in roughly that order.
A GenAI feature that works for a design partner is a genuine achievement and tells you almost nothing about whether the product works at a thousand times the volume. The pilot validated that the capability is useful. Scale tests whether it is viable.
Unit economics become a product decision
At scale, inference cost per user is a pricing input, not an infrastructure line item. This is the difference from conventional software, where marginal cost per user rounds to zero and pricing is set entirely by value.
A GenAI product has a real, variable, usage-linked cost of goods sold. At pilot size that cost is a rounding error and nobody models it. At volume, with one customer’s power users making hundreds of calls a day each, it determines whether the account is profitable.
Three things worth knowing before the pricing conversation:
Cost distribution is not normal. A small share of users generates a large share of spend, and it is usually the customers you most want to keep. Model the cost of your heaviest decile, not your average user, because the average conceals the accounts that matter.
Context length drives cost more than call volume. Growing a retrieval window from five chunks to twenty multiplies token cost on every call without a proportionate gain in answer quality. Measure the cost of context additions the way you would measure the cost of any feature.
Not every step needs the frontier model. Classification, extraction, routing and formatting steps are frequently well served by smaller, cheaper, faster models. Deciding per step rather than per product is the single largest cost lever most teams have not pulled, and it also improves latency.
OWASP’s Top 10 for LLM Applications, 2025 edition, lists Unbounded Consumption (LLM10:2025) as a named risk. It is filed as a security risk, and it is also the mechanism by which a product’s margin disappears: without per-tenant and per-user rate and spend limits, a single integration mistake on a customer’s side becomes your cost.
Move Your GenAI Product Beyond the Pilot
Address performance, reliability, cost, and integration challenges as your AI product grows from early users to enterprise-wide adoption.
Evaluation stops being a project and becomes infrastructure
At scale, evaluation has to run on every change, automatically, against a maintained dataset — the same way a test suite does. Manual review of a sample of outputs is a pilot-stage practice that does not survive a release cadence.
The reason is the surface area. A change to a prompt, a retrieval parameter, a chunking strategy, a model version or a tool schema can shift behavior on cases nobody thought to check. With ten users, someone notices. With ten thousand, the first signal is a support ticket from your largest customer.
What an evaluation harness needs to be useful in production:
– A dataset that grows from real traffic, including the cases that failed. Synthetic evaluation sets drift away from production usage, and the drift is invisible until something regresses.
– Graded outcomes, not binary pass/fail, for tasks where several answers are acceptable.
– Segment-level results, because aggregate accuracy hides the fact that the system got worse for one customer’s document format.
– A regression gate in the release pipeline, so a change that degrades a segment is caught before deployment rather than after.
The uncomfortable part is that building this costs real engineering time and produces no user-visible feature. Teams defer it until the first bad release, which is also the point at which they discover that reconstructing what changed is difficult, because prompts and parameters were not versioned.
Version prompts, retrieval configuration and model selection as code, in the same repository as the application, reviewed the same way. Once they live in a database that someone edits through an admin panel, the ability to answer “what changed?” is gone.
Models are deprecated on someone else’s schedule
Build on the assumption that every model you use will be withdrawn or changed, and that you will not choose when.
Provider model lifecycles are not aligned to your release calendar. A version gets deprecated, a default changes, pricing shifts, or behavior moves after a provider-side update. Products with a model identifier hard-coded across a dozen call sites discover this as a migration project with a deadline set externally.
The mitigations are ordinary software engineering:
Abstract the provider behind an interface. One place where model selection happens, configurable per step.
Keep a switching test. The evaluation harness above is what makes a model migration a measurable exercise rather than a leap. Without it, changing models is a decision made on vibes and discovered in production.
Instrument per-model behavior. When output quality drops, the first question is whether anything changed on the provider side. That is only answerable if per-model metrics are being recorded continuously.
This connects directly to the layer model in the anatomy of a production AI agent: abstracting the model is a layer-one decision whose cost is paid, or not paid, at layer five.
Tenant isolation is a retrieval problem
In a multi-tenant GenAI product, the retrieval layer is where data leaks between customers, and it does not announce itself.
Conventional multi-tenancy is well understood: row-level security, separate schemas, partition keys. Retrieval introduces a path around all of it. Embeddings are stored in a vector index. If the index is shared and the tenant filter is applied as a post-retrieval step rather than a pre-filter, a query can retrieve another tenant’s content and then discard it — except that the content has already reached the application, and any bug between those two points is a disclosure.
Design rules that hold:
1. Filter before retrieval, not after. The tenant boundary is a query constraint, not a result filter.
2. Respect user-level permissions inside a tenant. A document the requesting user cannot open in the source system must not be retrievable through the assistant. This is more common than cross-tenant leakage and just as serious.
3. Do not put one tenant’s content into a shared cache key. Semantic caching is an effective cost lever and a straightforward way to serve one customer’s answer to another.
4. Handle deletion. When a customer deletes a document, or leaves, the embeddings have to go too. Retention policy applies to derived data, which is the part most commonly missed.
OWASP lists Sensitive Information Disclosure (LLM02:2025) and Vector and Embedding Weaknesses (LLM08:2025) separately for this reason — the storage layer has its own failure modes.
Latency budgets do not survive naive retries
Set an end-to-end latency budget per user-facing action and design the failure path inside it. Retry logic borrowed from conventional services multiplies GenAI latency rather than smoothing it.
A conventional API call that fails is retried in milliseconds. A model call that times out at eight seconds and is retried costs sixteen, and a second retry costs twenty-four — by which point the user has left or the surrounding request has timed out. Under load, retries also arrive precisely when the provider is slowest, which is how a degradation becomes an outage.
What works at scale:
– A budget per action, stated, with the model call allocated a share of it — not the whole thing.
– Degrade rather than retry. A smaller, faster model, a cached answer, or a plainly labelled partial response beats a second attempt.
– Stream where the interaction allows it. Time to first token is what users experience as responsiveness; total generation time frequently is not.
– Fail visibly and specifically. “The assistant could not complete this — here is the underlying data” retains more trust than a spinner followed by an error.
Feedback has to be designed in, or it does not exist
The product needs a mechanism that captures which outputs were wrong and why, attached to the exact inputs, prompt version and model that produced them. Retrofitting this is expensive; designing it in costs very little.
Thumbs up and down alone is close to worthless — it records sentiment without diagnosis. What makes feedback usable for evaluation:
– The full input context, retrieved documents and prompt version stored against the interaction.
– A reason taxonomy short enough that people use it: wrong facts, missed context, wrong format, too slow, unsafe.
– Corrected outputs where users naturally produce them — an edited draft is the highest-quality training and evaluation signal a product generates.
– A path from flagged interaction into the evaluation dataset that does not require an engineer.
This closes the loop with the evaluation section: the harness is only as good as its dataset, and the dataset comes from production or it comes from imagination.
Build for Production, Not Just the Prototype
Develop GenAI applications with production-focused engineering across data, model integration, monitoring, security, and ongoing improvement.
What to build before scaling, not after
A defensible order of work for a team taking a validated GenAI feature toward volume:
| Priority | Capability | Why this order |
|---|---|---|
| 1 | Prompt, config and model versioning in code | Everything else is unanswerable without it |
| 2 | Per-tenant and per-user rate and spend limits | Caps the blast radius of both cost and abuse |
| 3 | Evaluation dataset from real traffic, with a regression gate | Makes every subsequent change measurable |
| 4 | Pre-filtered, permission-aware retrieval | Isolation is not retrofittable safely |
| 5 | Per-step model selection | The largest cost and latency lever |
| 6 | Structured feedback capture | Feeds 3, compounding over time |
Items 1 and 2 are the cheapest and prevent most of the expensive failures. Item 3 is the one that separates products that improve from products that drift.
For how these systems are built and taken to production, see generative AI development and AI application development built for production. For connecting them to systems of record without losing the audit trail, see AI integration services.
Frequently asked questions
What changes when a GenAI product scales?
Five things that did not exist at pilot size: inference cost becomes a pricing input, evaluation has to run automatically on every release, provider model deprecation becomes a recurring migration, tenant isolation has to be enforced in the retrieval layer, and latency budgets stop tolerating retries.
How do you control the cost of a generative AI product?
Set per-tenant and per-user rate and spend limits, choose models per step rather than per product, treat context length as a cost decision, and model the heaviest decile of users rather than the average. OWASP lists unbounded consumption as a named LLM risk for the same reason.
Do you need an evaluation harness for a GenAI product?
At scale, yes. Manual review of sampled outputs does not survive a release cadence, because a change to a prompt, retrieval parameter or model version can shift behavior on cases nobody checked. The harness needs a dataset drawn from real traffic and a regression gate in the release pipeline.
How do you prevent data leaking between tenants in a GenAI product?
Apply the tenant boundary as a pre-retrieval query constraint rather than a post-retrieval filter, respect user-level source permissions inside each tenant, keep tenant content out of shared cache keys, and delete embeddings when source documents are deleted.
Should GenAI products retry failed model calls?
Usually not. A retry on an eight-second timeout doubles the latency and arrives when the provider is already degraded. Degrading to a smaller model, a cached answer or a clearly labelled partial response preserves the latency budget and the user's trust.
How should prompts and model configuration be managed?
As code, in the application repository, versioned and reviewed like any other change. Once prompts live in a database edited through an admin panel, the question "what changed before this regression?" stops being answerable.
Build GenAI Products Ready to Scale
Start Your AI Project