Runtime boundary
Codapult keeps provider calls behind one server-side gateway. Product features call the gateway or the full pipeline; they do not create provider clients in route handlers or components.
Client UI / server action
│
▼
API route: auth → rate limit → Zod request validation → organization quota
│
▼
AI pipeline: budget → RAG context → input guardrails → tools → gateway
│
▼
Model catalogue → provider adapter → Vercel AI SDK
│
├── response / stream
├── output guardrails
└── usage, cost, latency, and request log
The route owns access control and request shape. The pipeline owns product policy. The gateway owns model selection, retries, fallback, AI SDK calls, usage normalization, and cost calculation.
Agents and batch jobs use the same pipeline as chat. Agents resolve their versioned prompt and organization-scoped tools before entering it; batch resolves its prompt version once and runs each item through the pipeline. API entry points apply the user rate limit and organization AI quota before work is scheduled.
End-to-end request flow
The same runtime boundary applies the important controls in a deliberate order. The exact path for an interactive request is:
request
→ auth / organization scope
→ quota and budget checks
→ model allow-list and routing
→ input guardrails
→ tools and optional RAG context
→ model generation or stream
→ output evaluation
→ usage and cost recording
→ organization-scoped request log
| Stage | What it does | Current implementation |
|---|---|---|
| Auth / organization scope | Confirms the session and active organization before the AI route proceeds. | src/app/api/ai/chat/route.ts, PipelineContext |
| Quota and budget | Applies per-user rate limits, plan credits, and optional daily or monthly AI budgets. | checkOrgQuota(), checkBudget() |
| Model routing | Rejects models outside the catalogue or allow-list, then applies retries and configured fallbacks. | src/lib/ai/models.ts, src/lib/ai/gateway/router.ts |
| Guardrails | Evaluates organization-scoped input and output rules; rules can block, warn, redact, or log. | src/lib/ai/guardrails/ |
| Tools and RAG | Resolves enabled organization-scoped tools and optionally injects tenant-scoped retrieved context. | src/lib/ai/tools/, src/lib/ai/rag.ts |
| Generation / stream | Uses the same gateway for generateText() and streamText(), with bounded tool steps. | src/lib/ai/gateway/router.ts |
| Usage and cost | Normalizes token usage, calculates provider cost, and records latency. | src/lib/ai/gateway/cost.ts, src/lib/ai/metering/ |
| Request log | Persists model, provider, prompt, organization, status, metadata, usage, and cost for inspection. | aiRequestLog, logRequest() |
Tool executions also emit compact activity events (ai.tool.executed or ai.tool.failed) with the organization, agent, tool, handler type, and duration. Arguments and results are deliberately excluded from these events; sensitive tools should add their own domain audit event when the business action itself must be reconstructable.
This is a shared execution path, not a promise that every AI feature has identical UX. Chat streams responses, agents can run multi-step tools, and batch jobs process asynchronously; they converge on the gateway and its provider, policy, and metering boundaries.
Runtime guarantees and boundaries
- Prompt versioning: prompts are stored with versions and the selected prompt ID can be attached to a gateway request and its usage log. Agent execution resolves the current configured version before calling the gateway.
- Tool permissions: tools are resolved only when enabled and available to the organization. Sensitive tools still need authorization inside their handler; registering a tool is not a substitute for domain permission checks.
- Tenant isolation: RAG retrieval, guardrail loading, tool lookup, request logs, quotas, and agent lookup accept the active organization scope. Do not derive that scope from arbitrary browser input.
- Fallback semantics: retries happen within a model attempt; the configured fallback chain is considered after a classified error, rate limit, or timeout, and every fallback model must pass the configured allow-list.
- Evaluation and testing: deterministic evals cover request validation, model allow-lists, guardrail behavior, and cost accounting. Provider-backed quality evals remain application-specific and should be versioned with their datasets.
- PII redaction: non-streaming responses can be blocked or redacted before they are returned. Streaming output is evaluated after completion because already-delivered tokens cannot be retracted. The gateway accepts
requireCompleteResponse: trueto force the complete-response path even if a caller also requests streaming; configure output guardrails on that route when pre-delivery redaction is required. - Streaming and cancellation: streaming is implemented through the AI SDK and the chat route now propagates the request
AbortSignalto the provider call. Batch cancellation still changes the persisted job state; it is not the same as cancelling every already-started batch item. - Audit trail: request logs include a compact summary of agent tool calls. The current
aiRequestLogis still an operational log, not a complete immutable audit trail of every tool side effect; add domain audit events around sensitive handlers when that evidence is required.
Capability coverage
This table is the contract boundary for the current AI foundation. “Application-specific” means the extension point is present, but the buyer must provide the product policy or domain implementation.
| Capability | Current coverage | Boundary to account for |
|---|---|---|
| Unified gateway for chat, agents, batch, and RAG | Implemented through the shared gateway pipeline. | The user experience differs by surface: chat streams, agents run tool steps, and batch runs asynchronously. |
| Shared guardrails and metering | Implemented for the full pipeline used by chat, agents, and batch. | Direct low-level provider calls or gatewayGenerate() calls bypass pipeline policy and should remain internal. |
| Tool permissions | Organization-scoped enabled tools are resolved before execution. | Tool handlers must enforce record-level and business authorization; registry access is not RBAC. |
| Tenant isolation | Organization scope is carried through RAG, guardrails, tools, agents, quotas, and logs. | Never accept the active organization from arbitrary browser input; validate membership at the route boundary. |
| Prompt versioning | Agents and batch resolve a prompt version and record its ID with the request. Gateway requests can attach promptId. | The built-in chat request does not currently expose prompt selection; product-critical chat prompts should be wired explicitly by the integrating route. |
| Fallback semantics | Classified provider errors, rate limits, and timeouts can move through the configured fallback chain. | Fallbacks still need compatible capabilities and valid credentials; they do not guarantee semantic equivalence between models. |
| Retries | Bounded retries are applied within a model attempt. | Retry limits and backoff are configuration; non-idempotent tool side effects need their own idempotency protection. |
| Model allow-list | Unknown models and configured disallowed models are rejected before generation. | Adding a provider requires a catalogue entry, adapter configuration, and pricing/capability data. |
| Cost calculation | Usage is normalized and provider cost is recorded on the gateway path. | Pricing metadata must be kept current when models or providers are added. |
| Quotas | Chat, gateway, and batch entry points apply the organization aiChat request quota; the pipeline also checks optional spend budgets. | Internal direct calls are not an independent quota boundary. Request quotas, token cost, and spend budgets are separate controls. |
| Evaluation and testing | Deterministic evals cover schema validation, model policy, guardrails, and cost accounting. | Provider-backed quality evals require buyer-owned datasets, acceptance criteria, and model-cost decisions. |
| Audit trail | Request logs and compact tool activity events capture operational evidence. | This is not an immutable business audit ledger. Sensitive handlers must emit domain audit events for reconstructable side effects. |
| Streaming and cancellation | Chat streaming propagates AbortSignal; complete responses support pre-delivery output checks. | Batch cancellation changes job state but does not guarantee cancellation of every provider call already in progress. |
| PII redaction | Input policies run before generation; complete responses can block or redact before delivery. | Streaming output is evaluated after tokens are delivered. Use requireCompleteResponse: true plus a configured output rule for strict pre-delivery handling. |
Source map
| Responsibility | Source |
|---|---|
| Public model catalogue and capabilities | src/lib/ai/models.ts |
| Provider and SDK model resolution | src/lib/ai/gateway/providers.ts |
| Retry, fallback, generate, and stream | src/lib/ai/gateway/router.ts |
| Budget, RAG, guardrails, tools, and logging orchestration | src/lib/ai/gateway/pipeline.ts |
| Token pricing and cost calculation | src/lib/ai/gateway/cost.ts and src/lib/ai/metering/ |
| Typed request boundary | src/lib/ai/api-schemas.ts |
| Streaming chat endpoint | src/app/api/ai/chat/route.ts |
| Persistent conversations | src/lib/ai/conversations.ts and src/app/api/ai/chat/conversations/ |
| Tools and agent execution | src/lib/ai/tools/ |
| Embeddings and retrieval | src/lib/ai/embeddings/, src/lib/ai/vector-store/, src/lib/ai/rag.ts |
| Prompt versions and A/B assignment | src/lib/ai/prompts/ |
| Deterministic runtime evals | src/lib/ai/evals/ |
Generate and stream paths
Use runPipeline for a complete non-streaming request. It performs the budget check, loads organization guardrails, optionally adds RAG context, resolves organization tools, calls the gateway, evaluates the output, and records usage.
Use runStreamPipeline for interactive chat. It performs the request-side policy checks before streaming and runs output evaluation in the completion callback for persisted analytics.
Streaming cannot retract tokens already delivered to a client. If output policy must block or redact content before the user sees it, send requireCompleteResponse: true together with the selected output guardrails. The gateway then uses the non-streaming path. A server-owned sensitive route should set this policy itself rather than relying on browser code to choose it. Buffering is an application-specific alternative when a delayed streaming-like UX is required.
{
"messages": [{ "role": "user", "content": "Summarize this record" }],
"stream": true,
"requireCompleteResponse": true,
"guardrails": ["customer-output-pii"]
}
requireCompleteResponse controls delivery timing; it does not create a guardrail. The referenced organization-scoped output rule must exist and use the desired block or redact action.
Extension rules
- Add a model to
src/lib/ai/models.tsbefore using it in configuration, UI, or pricing. - Add provider credentials to the server environment schema; never send them to the client.
- Add domain tools through the tool registry with a Zod-compatible parameter schema and an explicit organization scope.
- Keep tenant context in
PipelineContext; do not resolve organization ownership from browser-provided data alone. - Keep prompts versioned when the response is product-critical, and record the prompt ID with usage logs.
- Bound
maxSteps,maxTokens, request size, RAG chunks, retries, and budgets. - Treat model output and tool arguments as untrusted input. Apply validation, authorization, and guardrails before side effects.
Runtime evals
The built-in evals are deterministic contract checks around the provider boundary. They do not call an external model or spend tokens. They cover:
- request schema rejection;
- model catalogue and allow-list behavior;
- input blocking and output PII redaction;
- additive token-cost accounting.
Run them independently from the full suite:
pnpm test:ai-evals
Provider-backed quality evals belong to the product that owns the prompts and acceptance criteria. Keep those datasets versioned, redact user data, pin the model and temperature, and record latency, token usage, cost, and pass/fail reasons. Do not make external model calls part of the default unit-test suite.