Opticat item search MCP reviewPhase 1 discovery, validation, and path forward
Evaluation

How we prove the assistant’s answers can be trusted

The evaluation foundation is in place. Release proof is not complete yet. This page shows what we have, what is missing, and the test that closes the gap.

01

What evidence do we have today?

We have useful test material and a clear test design. That is different from having a completed, current release score.

Inputs ready to useGood starting point
50Historical QA cases

Real prior questions and observed outcomes. Diagnostic seed; not a current accuracy score.

100Scenario portfolio

Broad customer-journey coverage. A demonstration library—not 100 adjudicated acceptance cases.

Proof still being completedNot a release score
32Tool-shape cases

Eight tools × positive, negative, ambiguous, and partial/unavailable behavior. Preview validation is pending.

16Core comparison cases

Executive-facing regressions with guarded controls and pending current validation.

Current conclusionAs of 2026-08-24

Most tools and core cases still need one controlled current run before an executive can treat the result as release proof.

Eight task-level tools8 total
1 live verified2 live guarded5 need controlled retesting
Core customer cases16 total
0 implemented2 guarded14 need current validation
The 16 core comparison cases, one status each
CaseStatusWhy it holds this status
CORE-VIN-01Needs current validationVIN decode repaired after the live HVAC-continuation failure; preview retest pending.
CORE-VIN-02Needs current validationShares the repaired VIN path with CORE-VIN-01; preview retest pending.
CORE-XR-01Needs current validationReverse interchange recovery was added after the live Kia zero-result run.
CORE-XR-02Needs current validationPositive baseline retained; needs a rerun on the repaired interchange path.
CORE-FV-01Needs current validationWhole-token product-line resolution was repaired after the live GMB failure.
CORE-XR-03Needs current validationTyped interchange output needs a live completeness check against the Akebono set.
CORE-SS-01Needs current validationLifecycle supersession routing was repaired; a domain-approved chain is still required.
CORE-SS-02Needs current validationExplicit replaces/replaced-by labeling needs a live VW-chain confirmation.
CORE-SS-03Needs current validationInterchange-vs-supersession labeling needs a live GM rerun.
CORE-PART-01Needs current validationBounded fitment and budget behavior need a live multi-turn replay.
CORE-VEH-01Needs current validationJob-completion coverage beyond pumps and gasket is not yet demonstrated.
CORE-VEH-02Needs current validationBrand-completeness disclosure needs a live rerun of the novice flow.
CORE-PD-01Needs current validationAttribute unit fidelity needs confirmation in the repaired details UI.
CORE-PD-02Needs current validationDirect asset delivery from part details needs a live image render check.
CORE-REL-01GuardedClosed-world completeness states are enforced; guarded pending broader replay.
CORE-SYM-01GuardedSymptom-side recommendations stay behind the qualifier guardrail.

Do not report 26% accuracy. The 13 historical “pass” labels came from different test moments. They were not scored against one frozen system and one approved rulebook.

02

Who decides what the right answer is?

Engineering can automate retrieval and checks. OptiCat domain experts must define catalog truth, material qualifiers, relationship meaning, and the acceptance threshold.

The evaluation criteria are a shared product artifact—not an AI guess.

The MCP can parse wording, fetch quickly, preserve evidence, and enforce approved rules. It cannot decide on its own whether a catalog relationship is commercially equivalent, which qualifier changes safe fitment, or which production risk OptiCat will accept.

OptiCat catalog/API owner
Owns
Catalog truth, operation semantics, key scope, rate limits, and data drift.
Decides
Whether a record is reachable, missing, not entitled, or represented differently upstream.
Signs off
Raw API ground truth and entitlement profile
Fitment/category specialists
Owns
Part-category rules, material qualifiers, relationship meaning, and safe fitment interpretation.
Decides
Which engine, trim, side, position, quantity, note, or lifecycle field changes the answer.
Signs off
Case expectations and disputed fitment/relationship outcomes
OptiCat product owner
Owns
Customer journeys, useful response shape, clarification tolerance, and demonstration scope.
Decides
What the answer must include and when asking, qualifying, or abstaining is acceptable.
Signs off
Case contract and supported conference journeys
Tenexity MCP engineer
Owns
Tool contracts, API orchestration, completeness, typed errors, and evidence preservation.
Decides
How domain rules become deterministic service behavior across hosts.
Signs off
MCP fidelity and contract tests
Tenexity evaluation engineer
Owns
Frozen runs, deterministic checks, trace capture, grading workflow, and KPI computation.
Decides
Where a failure entered the API, MCP, routing, or final-answer layer.
Signs off
Reproducible result and failure localization
Executive approver
Owns
Production thresholds, risk tolerance, investment, schedule, and final acceptance.
Decides
Whether evidence is sufficient to authorize a demonstration build or production step.
Signs off
Release gate and commercial authorization
Gate
Case lifecycle from draft to release decision

Automation handles repeatability; named experts handle domain judgment and disputed truth.

  1. 01
    Draft the caseEvaluation engineer

    Capture the question variants, intended journey, required evidence, and prohibited claims.

  2. 02
    Approve domain truthOptiCat experts

    Confirm expected answers, material qualifiers, relationship meaning, and entitlement assumptions.

  3. 03
    Freeze the systemMCP + evaluation

    Record code, deployment, model, prompt, tools, key scope, settings, and catalog version.

  4. 04
    Capture every layerEvaluation engineer

    Retain redacted API, MCP, routing, customer answer, status, and latency evidence.

  5. 05
    Run deterministic checksEvaluation automation

    Check existence, exact fitment, qualifiers, totals, errors, traceability, and prohibited claims.

  6. 06
    Adjudicate exceptionsOptiCat domain reviewer

    Resolve data drift, ambiguous expectations, and disagreements without silently changing the rubric.

  7. 07
    Publish the decisionExecutive approver

    Report safety gates and journey-level KPIs, then accept, repair, or block release.

03

How do we test an answer?

Start with the right answer, keep the evidence from every handoff, and name the first place the result went wrong.

  1. Step 01
    Set the expected answer

    Agree on what a good answer must include, when the assistant should ask a question, and what it must never claim.

  2. Step 02
    Confirm the catalog truth

    Check the current OptiCat record first, including access limits, qualifiers, totals, and missing data.

  3. Step 03
    Compare the full answer chain

    Compare the catalog result, tool response, tool choices, and customer-facing answer to find where a problem began.

  4. Step 04
    Record the decision and owner

    Give the run one clear grade, name the first failed layer, and assign the work needed before release.

Rules for a fair testFive controls that keep the result honest and repeatable
  1. 01
    Freeze the system being tested

    Record the repository SHA, deployed MCP version, host instructions, model settings, API-key scope, catalog version, and timestamp so a rerun means the same thing.

  2. 02
    Write the answer contract first

    Every case states the intended journey, required evidence, acceptable clarification, facts that must appear, and claims the assistant must never make.

  3. 03
    Test the same meaning in different language

    Use catalog-style wording, normal customer language, shorthand, misspellings, follow-ups, and incomplete requests to prove that routing is based on meaning—not a memorized phrase.

  4. 04
    Compare every layer

    Keep raw API evidence, the MCP result, tool arguments, the final answer, and the user-visible state. That shows exactly where a failure entered the chain.

  5. 05
    Adjudicate instead of guessing

    When the API, OptiCat Online, historical notes, or two reviewers disagree, a named catalog owner resolves the expected answer and the disagreement remains visible.

Evidence kept at every layerSix records that show where a wrong answer entered the chain
  1. 01
    Case contractQuestion variants + expected intent

    We know what success, clarification, and prohibited claims mean before running the system.

  2. 02
    Upstream truthRaw OptiCat response + entitlement

    The relevant record is reachable, missing, scope-blocked, or different from historical ground truth.

  3. 03
    MCP fidelityStructured tool result

    Identity, totals, qualifiers, relationships, lifecycle, assets, and errors survived the server boundary.

  4. 04
    Routing and judgmentIntent, entities, tool calls, choices

    The assistant chose the right journey, rerouted safely, or asked only for a detail that changes the answer.

  5. 05
    Customer answerFinal response + product/vehicle cards

    Every claim is supported, useful, complete enough, and clear about assumptions or limitations.

  6. 06
    Decision recordGrade + owner + failure layer

    The result is reproducible, reviewable, and actionable for release or repair.

How failures are labeledOne grade that tells engineering what kind of repair is needed
Correct

The answer is supported and satisfies the case contract.

grounded_correct
Correct but incomplete

The returned facts are supported, but required catalog coverage or detail is missing.

grounded_incomplete
False negative

The answer denies a record that the entitled, complete upstream evidence contains.

false_negative
Unsupported positive

The answer adds a part, fitment, attribute, or relationship that evidence does not support.

hallucinated_positive
Contradiction

The answer conflicts with its own earlier result or already-returned evidence.

contradiction
Qualifier error

The part may exist, but engine, trim, side, position, quantity, or another condition is wrong.

qualifier_error
Infrastructure failure

The run cannot be judged because a service, credential, timeout, or tool boundary failed.

infra_failure
Full test procedureSeven steps from approved expectation to published decision
  1. 01
    Approve the case contract

    OptiCat confirms the intended journey, must-return facts, acceptable clarifications, prohibited claims, and required catalog fields.

  2. 02
    Freeze and identify the run

    Record code, deployment, model, prompt, key scope, catalog version, settings, and a fresh conversation ID.

  3. 03
    Establish upstream truth

    Run the required OptiCat operations to completeness and classify the record as reachable, missing, scope-blocked, or ground-truth drift.

  4. 04
    Run MCP and customer paths

    Execute the direct tool path and full assistant path, retaining redacted arguments, structured results, final answer, errors, and latency.

  5. 05
    Repeat meaning variants

    Replay traditional catalog wording, natural language, shorthand, misspellings, and follow-ups while keeping the intended meaning fixed.

  6. 06
    Score and localize

    Apply deterministic checks first, then human review where needed; assign exactly one grade and the first layer where the failure appeared.

  7. 07
    Publish the decision

    Report safety gates, accuracy by journey, infrastructure failures, variant consistency, and disputed cases—never only one blended percentage.

04

Will it work when customers say it differently?

The same request should reach the same safe answer whether the customer uses catalog terms, normal language, shorthand, or an incomplete question.

Customer-language example

Vehicle → parts

What the customer asked
I need front brake pads for a ’97 Mustang GT.
What the system should understand1997 Ford Mustang GT; Disc Brake Pad Set; Front
  • Expand ’97 to 1997
  • Translate ‘brake pads’ into catalog part-type candidates
  • Recognize GT as a material submodel qualifier
When it must ask a question

Ask about position, trim, or engine only when the returned catalog choices produce different valid parts.

What it must never claim

Do not mix front and rear, combine Base/GT/Cobra results, or present the first page as the whole catalog.

Catalog path

Year → make → model → submodel/engine → part type → position

System route · search_parts_for_vehicle
05

What must be true before release?

Safety is a hard gate. Quality and speed are reported by customer journey so one easy test cannot hide a dangerous failure.

SOW success metricsSeven named measures, one frozen baseline still to complete

Targets are recommendations until OptiCat domain and executive owners approve them against the current replay.

Retrieval accuracy

Adjudicated cases where the system retrieves the correct, sufficiently complete catalog evidence for the intended journey.

OptiCat domain + Tenexity evaluation
Target · ≥95% by journey after baselineMeasurement in progress
Hallucination rate

Answers containing a part, fitment, attribute, specification, or relationship claim absent from current evidence.

Tenexity MCP/evaluation
Target · 0%Measurement in progress
Abstention rate

Insufficient-evidence cases that qualify or abstain rather than fabricate a catalog fact.

OptiCat product + Tenexity evaluation
Target · ≥99% safe handlingMeasurement in progress
Clarification rate

Material ambiguities resolved by one focused question without unnecessary customer friction.

OptiCat product/domain
Target · ≥90% effectiveMeasurement in progress
Response latency

End-to-end time from customer question to the final supported response, reported by journey.

Tenexity engineering
Target · p95 ≤15s without weakening checksMeasurement in progress
Citation coverage

Customer-facing catalog claims linked to current API/MCP evidence retained for the run.

Tenexity MCP/evaluation
Target · 100%Measurement in progress
Incorrect part recommendation rate

Recommended products that fail exact identity, lifecycle, or requested-vehicle fitment validation.

OptiCat domain + Tenexity evaluation
Target · 0%Measurement in progress
Safety gates

A release stops if the assistant invents catalog facts or recommends the wrong part.

2 measures
MeasureTarget
Unsupported catalog claimsAnswers containing a part, specification, relationship, or fitment claim absent from current evidence.0%
Incorrect part recommendationsRecommended products that fail existence or exact-fitment validation.Target 0%
Evidence requirements

Every firm claim must have current proof, and every limited result must say it is incomplete.

3 measures
MeasureTarget
Fitment verification coverageUnconditional ‘fits’ claims with exact application evidence.100%
Evidence coveragePart, specification, and fitment claims linked to current source evidence.100%
Truncation disclosureLimited responses that clearly show total, returned count, and incomplete state.100%
Answer quality

The assistant must find the right answer, handle uncertainty, and keep the same meaning across normal customer wording.

4 measures
MeasureTarget
Grounded retrieval accuracyAdjudicated grounded-correct or grounded-incomplete cases, reported overall and by journey.≥95% after baseline
Clarification effectivenessClarifying questions that resolve a scored ambiguity on the next turn.≥90%
Safe clarification / abstentionInsufficient-evidence cases that ask, qualify, or abstain instead of guessing.≥99%
Meaning-preserving variantsEquivalent phrasings that reach the same journey and an equivalent evidence-backed disposition.Baseline, then gate by cohort
Operations

A safe answer still needs to arrive fast enough for a normal customer conversation.

1 measure
MeasureTarget
Normal-flow latencyEnd-to-end time from customer question to final response without weakening verification.p95 ≤15s
What happens next

Four steps to a release decision

Each step closes a specific evidence gap. None can be replaced by a polished demo.

  1. 01
    Approve what good looks like

    OptiCat confirms the expected answers, acceptable questions, required evidence, and claims that are never allowed.

  2. 02
    Freeze one test system

    Record the code, deployment, model, instructions, catalog access, and settings so every result can be repeated.

  3. 03
    Run and review the evidence

    Test the direct tools and customer conversations, then resolve any disagreement about the right answer.

  4. 04
    Publish the release decision

    Report the safety gates, results by customer journey, remaining failures, and the owner of each fix.

Release gateShip only when the evidence supports the answer.

That means zero unsupported recommendations, verified tool shapes, completed current testing, resolved ground-truth disputes, and a named owner for every remaining failure.

Review the core cases
Continue the storyBefore and after architecture
View Demo