Capability profile · AI and Web Applications
AI features with evidence and a review path
A capability profile for integrating retrieval and model output into an application without hiding sources, uncertainty, or human review.

01 Scope the decision
Identify the user, the allowed inputs, the consequence of an incorrect suggestion, and the cases that require a person.
02 Retrieve with access rules
Find only material the current user may see and retain identifiers that let the interface show the source.
03 Constrain the output
Validate model responses against a typed contract; handle refusal, timeout, and malformed output explicitly.
04 Evaluate and improve
Keep representative test cases, capture corrections, and compare quality and cost before changing a model or prompt.
Capability profile — not a client case study. This is an illustrative architecture for AI application work. No customer deployment, accuracy score, or business result is being claimed.
Put the model inside a product boundary
A model call is not a workflow. A useful application needs to decide which data may be supplied, what the model is allowed to return, and what happens when the answer is uncertain or unavailable. The application—not the prompt—must enforce authorization and validate the response.
Keep evidence attached
For retrieval-based features, preserve document identifiers, versions, and the specific passages used. Show those sources beside the generated suggestion so a reviewer can verify them. Retrieval must apply the same access rules as the underlying records; hiding an unauthorized result after generation is too late.
Make review a first-class state
Represent an answer as a suggestion, not an irreversible fact. Let a user accept, edit, reject, or escalate it. Store the model and prompt version along with the minimum audit data needed to investigate behavior. Avoid retaining raw prompts or sensitive documents longer than the product requires.
Test the system, not just the prompt
- Use a small, representative evaluation set that includes ambiguous and out-of-scope requests.
- Measure retrieval quality separately from answer quality.
- Test authorization boundaries, timeouts, rate limits, and malformed output.
- Track latency and cost per workflow, with explicit limits on retries.
Start with the narrowest task where a human can check the result. If the product cannot explain where its input came from or recover safely from a bad answer, adding a more capable model does not fix the underlying design.
Indicative technology options
These are representative choices, not a record of a deployed client project. The right stack depends on the product and its constraints.
- TypeScript or Python
- Model APIs or local inference
- Search and retrieval
- Application APIs
- Evaluation data