Real prior questions and observed outcomes. Diagnostic seed; not a current accuracy score.
Broad customer-journey coverage. A demonstration library—not 100 adjudicated acceptance cases.
The evaluation foundation is in place. Release proof is not complete yet. This page shows what we have, what is missing, and the test that closes the gap.
We have useful test material and a clear test design. That is different from having a completed, current release score.
Real prior questions and observed outcomes. Diagnostic seed; not a current accuracy score.
Broad customer-journey coverage. A demonstration library—not 100 adjudicated acceptance cases.
Eight tools × positive, negative, ambiguous, and partial/unavailable behavior. Preview validation is pending.
Executive-facing regressions with guarded controls and pending current validation.
Most tools and core cases still need one controlled current run before an executive can treat the result as release proof.
Do not report 26% accuracy. The 13 historical “pass” labels came from different test moments. They were not scored against one frozen system and one approved rulebook.
Engineering can automate retrieval and checks. OptiCat domain experts must define catalog truth, material qualifiers, relationship meaning, and the acceptance threshold.
The MCP can parse wording, fetch quickly, preserve evidence, and enforce approved rules. It cannot decide on its own whether a catalog relationship is commercially equivalent, which qualifier changes safe fitment, or which production risk OptiCat will accept.
Automation handles repeatability; named experts handle domain judgment and disputed truth.
Capture the question variants, intended journey, required evidence, and prohibited claims.
Confirm expected answers, material qualifiers, relationship meaning, and entitlement assumptions.
Record code, deployment, model, prompt, tools, key scope, settings, and catalog version.
Retain redacted API, MCP, routing, customer answer, status, and latency evidence.
Check existence, exact fitment, qualifiers, totals, errors, traceability, and prohibited claims.
Resolve data drift, ambiguous expectations, and disagreements without silently changing the rubric.
Report safety gates and journey-level KPIs, then accept, repair, or block release.
Start with the right answer, keep the evidence from every handoff, and name the first place the result went wrong.
Agree on what a good answer must include, when the assistant should ask a question, and what it must never claim.
Check the current OptiCat record first, including access limits, qualifiers, totals, and missing data.
Compare the catalog result, tool response, tool choices, and customer-facing answer to find where a problem began.
Give the run one clear grade, name the first failed layer, and assign the work needed before release.
Record the repository SHA, deployed MCP version, host instructions, model settings, API-key scope, catalog version, and timestamp so a rerun means the same thing.
Every case states the intended journey, required evidence, acceptable clarification, facts that must appear, and claims the assistant must never make.
Use catalog-style wording, normal customer language, shorthand, misspellings, follow-ups, and incomplete requests to prove that routing is based on meaning—not a memorized phrase.
Keep raw API evidence, the MCP result, tool arguments, the final answer, and the user-visible state. That shows exactly where a failure entered the chain.
When the API, OptiCat Online, historical notes, or two reviewers disagree, a named catalog owner resolves the expected answer and the disagreement remains visible.
We know what success, clarification, and prohibited claims mean before running the system.
The relevant record is reachable, missing, scope-blocked, or different from historical ground truth.
Identity, totals, qualifiers, relationships, lifecycle, assets, and errors survived the server boundary.
The assistant chose the right journey, rerouted safely, or asked only for a detail that changes the answer.
Every claim is supported, useful, complete enough, and clear about assumptions or limitations.
The result is reproducible, reviewable, and actionable for release or repair.
The answer is supported and satisfies the case contract.
grounded_correctThe returned facts are supported, but required catalog coverage or detail is missing.
grounded_incompleteThe answer denies a record that the entitled, complete upstream evidence contains.
false_negativeThe answer adds a part, fitment, attribute, or relationship that evidence does not support.
hallucinated_positiveThe answer conflicts with its own earlier result or already-returned evidence.
contradictionThe part may exist, but engine, trim, side, position, quantity, or another condition is wrong.
qualifier_errorThe run cannot be judged because a service, credential, timeout, or tool boundary failed.
infra_failureOptiCat confirms the intended journey, must-return facts, acceptable clarifications, prohibited claims, and required catalog fields.
Record code, deployment, model, prompt, key scope, catalog version, settings, and a fresh conversation ID.
Run the required OptiCat operations to completeness and classify the record as reachable, missing, scope-blocked, or ground-truth drift.
Execute the direct tool path and full assistant path, retaining redacted arguments, structured results, final answer, errors, and latency.
Replay traditional catalog wording, natural language, shorthand, misspellings, and follow-ups while keeping the intended meaning fixed.
Apply deterministic checks first, then human review where needed; assign exactly one grade and the first layer where the failure appeared.
Report safety gates, accuracy by journey, infrastructure failures, variant consistency, and disputed cases—never only one blended percentage.
The same request should reach the same safe answer whether the customer uses catalog terms, normal language, shorthand, or an incomplete question.
“I need front brake pads for a ’97 Mustang GT.”
Ask about position, trim, or engine only when the returned catalog choices produce different valid parts.
Do not mix front and rear, combine Base/GT/Cobra results, or present the first page as the whole catalog.
Safety is a hard gate. Quality and speed are reported by customer journey so one easy test cannot hide a dangerous failure.
Targets are recommendations until OptiCat domain and executive owners approve them against the current replay.
Adjudicated cases where the system retrieves the correct, sufficiently complete catalog evidence for the intended journey.
OptiCat domain + Tenexity evaluationAnswers containing a part, fitment, attribute, specification, or relationship claim absent from current evidence.
Tenexity MCP/evaluationInsufficient-evidence cases that qualify or abstain rather than fabricate a catalog fact.
OptiCat product + Tenexity evaluationMaterial ambiguities resolved by one focused question without unnecessary customer friction.
OptiCat product/domainEnd-to-end time from customer question to the final supported response, reported by journey.
Tenexity engineeringCustomer-facing catalog claims linked to current API/MCP evidence retained for the run.
Tenexity MCP/evaluationRecommended products that fail exact identity, lifecycle, or requested-vehicle fitment validation.
OptiCat domain + Tenexity evaluationA release stops if the assistant invents catalog facts or recommends the wrong part.
Every firm claim must have current proof, and every limited result must say it is incomplete.
The assistant must find the right answer, handle uncertainty, and keep the same meaning across normal customer wording.
A safe answer still needs to arrive fast enough for a normal customer conversation.
Each step closes a specific evidence gap. None can be replaced by a polished demo.
OptiCat confirms the expected answers, acceptable questions, required evidence, and claims that are never allowed.
Record the code, deployment, model, instructions, catalog access, and settings so every result can be repeated.
Test the direct tools and customer conversations, then resolve any disagreement about the right answer.
Report the safety gates, results by customer journey, remaining failures, and the owner of each fix.
That means zero unsupported recommendations, verified tool shapes, completed current testing, resolved ground-truth disputes, and a named owner for every remaining failure.