AI Infrastructure and Evaluation

Use this page for the cloud-neutral production envelope around an AI system and the evaluation loop that proves changes are safe. AWS service mappings belong in AWS AI services.

Part 6 of 7: AI Agents → Infrastructure and evaluation → AWS AI Services. This page integrates the earlier layers; it owns their security, reliability, operations, and release process rather than re-explaining their mechanics.

1. Production reference architecture

user → identity → API/application → authorization → AI orchestrator
                                                    ├─ model gateway
                                                    ├─ retrieval
                                                    ├─ tools/workflows
                                                    ├─ policy/guardrails
                                                    └─ state/cache
                         ← validated response + citations + audit trace ←

2. Ownership, data governance, and responsible AI

A company can run several maturity patterns at once; use the least complex one that safely serves the workload.

Pattern Use it when Add next
Bounded product integration One low-risk use case needs a model response. Evaluation, data classification, tracing, and cost limits before it becomes a production dependency.
Repeatable product pattern Several flows need the same prompts, model routes, retrieval, or policy. Shared adapters and standards, while the product still owns its user journey and domain workflow.
Shared AI platform Multiple teams need consistent identity propagation, approved model access, retrieval connectors, tool boundaries, telemetry, and evaluation. A paved road with reusable controls—not a central team that owns every feature.
Bounded agent extension Tool choice or sequence depends on runtime observations. The deterministic workflow, authorization, durable state, approvals, and budgets described on AI Agents.

Keep ownership explicit:

A prototype becomes unsafe when it reaches production without those owners and controls. Conversely, do not centralize every feature before a repeated need exists. Measure adoption by completed business tasks, quality, risk, latency, and cost—not model-call volume.

2.1 Data virtualization, ownership, and integrity

Data virtualization exposes a logical view across data sources without first requiring all their data to be copied into one central store. Consumers query that view while source systems remain responsible for their records. Implementations may still cache or materialize results. AWS’s data virtualization overview.

For an AI system, the architectural implication is that shared access still needs explicit ownership and integrity controls:

For example, a stock assistant can join warehouse and product views while each team owns its source records. An incorrect join can duplicate stock counts even when both sources are correct. Virtualization reduces unnecessary copies; it does not automatically repair data quality, transfer ownership, or provide consistent snapshots across independent systems. These are design responsibilities, consistent with AWS’s data ownership and governance architecture example.

2.2 Responsible AI dimensions

AWS describes eight dimensions, often called pillars in study material. The examples below translate them into controls for this architecture. AWS Responsible AI dimensions.

Dimension Practical question Example control
Fairness Who experiences worse outcomes? Evaluate error rates across relevant groups.
Explainability Can an output be understood and assessed? Provide validated feature-attribution information.
Privacy and security Is data used appropriately and protected? Minimize inputs and enforce access controls.
Safety Could outputs or actions cause harm? Test harmful requests and escalation paths.
Controllability Can people steer or stop the system? Human review, overrides, and rollback.
Veracity and robustness Does it remain correct under difficult inputs? Grounding checks and adversarial evaluation.
Governance Who owns decisions and oversight? Approval records, model documentation, and audits.
Transparency Do stakeholders know how AI is being used? Disclose purpose, limitations, and data practices.

Transparency concerns disclosure; explainability concerns understanding outputs. Publishing a model card improves transparency but does not by itself explain every prediction. A fluent generated explanation is not proof of the model’s actual decision process. AWS implementations include Model Cards and Clarify and Audit Manager.

3. Identity, authorization, and tenant isolation

4. Networking and service boundaries

5. Data protection and secrets

6. AI-specific security threats

Delimiters and a clear instruction hierarchy help the model distinguish instructions from data, but neither is an enforcement boundary. Treat all model-visible content as potentially hostile; enforce authorization, allowed destinations, and action policy in deterministic application and tool controls.

7. Scalability and quota management

8. Reliability and degraded modes

9. Latency engineering

serial request latency ≈ queueing + network + identity + retrieval
                       + reranking + context build + model prefill
                       + model decoding + tools

This is a serial-path approximation. Concurrent stages overlap, while agent loops, retries, and additional model calls add work to the critical path. Measure end-to-end latency directly; component p95 values cannot simply be added to obtain end-to-end p95.

10. Cost engineering

Start with model-call pricing, then add the costs outside that call: ingestion, embeddings, reranking, storage/indexing, tools, queues, logs, transfer, idle capacity, and human review. Attribute retries and agent loops to the originating task.

Use the latency measurements above to avoid paying for unnecessary work. Routing choices belong in Models; serving and caching choices follow below. Report total cost per successful task by tenant and task type so an expensive cohort is visible.

10.1 Serving and caching decision rules

Need Prefer Why
User waits for a short answer Synchronous inference, often streaming Immediate request/response path; streaming improves perceived responsiveness
Long-running, bursty, or non-interactive work Asynchronous job + queue Decouples callers from duration and absorbs load; needs durable status/idempotency
Large known dataset Batch inference Throughput/cost optimization where per-item immediate response is unnecessary
Identical safe request repeats Result cache Avoids an unnecessary model invocation
Semantically equivalent safe question repeats Semantic cache Matches meaning, not only identical bytes; evaluate false matches carefully
Stable common prompt prefix Prompt/prefix cache where the provider supports it Reuses repeated input processing; sensitive/tenant context must not cross boundaries

11. Observability

12. Evaluation architecture

change: prompt / model / parser / chunking / index / tool / policy
                                  ↓
versioned evaluation suite → compare baseline and slices → deploy or reject
                                  ↓
                      canary + production feedback

The evaluation suite composes specialist measurements; it does not duplicate them:

12.1 Golden datasets and reference answers

A golden dataset, also called a ground-truth or reference evaluation dataset, contains representative inputs with expert-validated expected answers or outcomes. For example, 200 customer questions paired with reviewed reference answers let a team compare prompt versions and detect regressions under the same conditions.

Dataset role What it is used for
Training data Learn or fine-tune model weights.
Validation data Select configurations, hyperparameters, or prompt variants during development.
Held-out test data Assess the chosen system on examples not used to tune it.
Golden/reference dataset Supply trusted expected results; it can support development evaluation or a separately held-out test suite.

“Golden” describes the quality of the reference, not a rule that evaluation happens only after deployment. “Test corpus” is a broader valid term, but does not specifically imply expert-reviewed question–answer pairs. Keep final test cases separate from tuning examples; version the references and rubric, cover difficult/no-answer cases, and review stale answers. For open-ended generation, score factual support and required content rather than demanding an exact wording match. See Bedrock evaluation datasets and ground-truth responses.

12.2 Precision, recall, F1, and accuracy

For binary classification, first define the positive class. If positive means a defective product, the confusion matrix is:

  Actually defective Actually acceptable
Predicted defective True positive (TP) False positive (FP): false alarm
Predicted acceptable False negative (FN): missed defect True negative (TN)
Metric Formula Question
Precision TP / (TP + FP) Of the flagged products, how many were defective?
Recall TP / (TP + FN) Of all defective products, how many did we catch?
F1 2 × precision × recall / (precision + recall) = 2TP / (2TP + FP + FN) How well do we balance precision and recall?
Accuracy (TP + TN) / (TP + TN + FP + FN) What fraction of all predictions were correct?

F1 is a harmonic mean, not an arithmetic average; it excludes true negatives. Accuracy includes true negatives and can hide missed positives when they are rare. Choose a metric according to the consequences of false alarms and missed cases. Google’s classification metrics.

Worked example: among 1,000 products, 100 are defective. The model flags 80: 60 are defective and 20 are acceptable. Thus TP=60, FP=20, FN=40, and TN=880:

A model that marks every product acceptable still gets 90% accuracy, while detecting no defects. Its recall is zero; precision has a zero denominator, so report the metric library’s undefined-value convention explicitly. Raising a decision threshold typically trades recall for precision; tune it on validation data. For multiclass problems, state whether the reported result is per-class, macro, micro, or weighted. scikit-learn classification metric conventions.

Retrieval Precision@K and Recall@K apply the same relevant-versus-selected idea to a ranked result set. Classification accuracy alone does not measure ranking, answer grounding, or generative quality.

13. Business and adaptability metrics

Metric What it answers
Efficiency Are compute, storage, energy, time, and human-review resources producing enough successful work for their cost?
Conversion rate What percentage of users complete the desired action, such as a purchase or sign-up?
User satisfaction Do users report that the experience and output are useful?
Cross-domain performance Does quality remain acceptable across distinct contexts such as healthcare, finance, and entertainment?

These measure different things. A system can be efficient but unpopular, or satisfy users while consuming too many resources.

14. LLM-as-judge and human evaluation

These are empirically observed limitations, not only theoretical concerns: the MT-Bench judge study examines position, verbosity, and self-enhancement biases. Agreement with another model is not a substitute for an appropriate human or reference standard.

15. Production evaluation and release gates

Use the business measures from section 13, the operational signals from observability, and the specialist RAG/agent measures linked in evaluation architecture. Add user correction, abandonment, and confirmed incidents to identify failures missed offline.

Regression gates should test the actual failure

For property descriptions that invented amenities and omitted square footage, keep historical inputs and verified facts in a versioned golden set. Assert the required square-footage value and compare structured amenity claims with the source’s allowed facts. Semantic similarity can remain high despite one critical invented fact. An exact string check also cannot detect every paraphrased hallucination; combine structured constraints with calibrated factuality evaluation.

Candidate prompt version → generate on held-out inputs
                         → required-field / source-fact assertions
                         → factuality + quality evaluation
                         → thresholds pass? → promote / reject

Run those gates before production promotion, then monitor/canary with rollback. Human creativity/persona comparisons need an expert rubric, and pre-deployment bias assessment needs the relevant protected-group metrics; a generic “quality” score does not answer every question. AWS evaluation settings and bias gates.

16. Troubleshoot the layer that failed

Bad result / failed task
        ↓
Was the needed evidence retrieved and authorized?
  No → parsing / chunking / query / filter / ranking problem
  Yes
        ↓
Did the prompt and selected model use the evidence correctly?
  No → context construction / prompt / model / decoding problem
  Yes
        ↓
Did an agent or tool choose/execute the correct action?
  No → schema / permission / workflow / tool-result problem
  Yes
        ↓
Did the platform meet the request contract?
  No → timeout / throttling / quota / retry / cache / deployment problem

Use the request trace to locate the first failed boundary, then add a regression case at that layer. Continue to AWS AI Services to implement these controls on AWS.

Contents