Klariton feature Test Center

Test your advisor before your customers do.

A test run sends questions through exactly the chat pipeline that serves your customers: typical questions, generated edge cases, eight security probes from a documented attack corpus, and your own test rules. Every answer then goes through hard and soft checks and through a panel of three independent models. What comes out is a traffic light and a list that starts with the red cases.

The test run

What gets tested is the real advisor, not a copy of it.

Every test question runs through the same pipeline as production, in test mode and without persistence. Every question gets a fresh session token, so the advisor answers it in isolation and carries no memory over from the previous one.

01Question set
Four sources, one set.

Typical questions, generated edge cases, eight security probes and your own rules. If the budget gets tight, the set is trimmed and the number of dropped questions is reported.

no silent trim
02Execution
Real pipeline, fresh token.

Every question is asked the way it would be in production. Nothing is persisted, and a failure on a single question does not abort the run, it is recorded as that question failing.

test mode
03Checks
Hard and soft, counted separately.

Answer present, evidence matching the source type, no inadmissible promise, no leak. Soft checks warn, hard checks decide whether a case passes.

per answer
04Judge panel
Three verdicts, one synthesis.

Three independent models rate the same answer against the same rubric. The synthesis is deterministic and needs no further model call.

3 judges
Security probes

75 documented attack patterns, eight classes, every run draws from every class.

The corpus comes out of two test matrices and covers eight OWASP LLM classes, from direct prompt injection through system prompt extraction, data leak and hallucination provocation to jailbreak, role play, scope overreach and lead abuse. A standard run deterministically takes one probe per class, so eight of them, so that every run checks security without firing the whole corpus.

Deterministic Same selection on every run.

The sample is not rolled at random, it takes the first probe of each class. That keeps two runs comparable instead of differing in what they picked.

8 of 75
Leak scan Against your real confidential values.

On security probes the answer is checked against the metafields of your catalogue that are marked confidential, such as purchase price, margin or supplier. A literal match is a hard failure, not a hint.

hard failure
Your own rules Your questions always run.

Test questions you add yourself are treated like the security probes: they are never trimmed away when the question budget of a run gets tight.

never trimmed
Scoring

Passed means: every hard check passed.

Every answer gets a list of individual checks, each with a severity and a readable reason. Soft checks produce a warning and turn the run amber. A case counts as passed only when all hard checks passed.

Judge panel

Three independent models rate the same answer against the same rubric: factual accuracy, whether the question was actually answered, tone for a buying conversation, no inadmissible promise, no data leak. The synthesis runs without a further model call: the majority decides, a tie goes to the stricter verdict, and dissent is stated in the reasoning.

Triggers

Manual, weekly, and whenever somebody gives a thumbs down.

A test run is not something you have to remember. It has a fixed rhythm, a button, and an automatic trigger coming out of live operation.

Manual At the push of a button in the Studio.

You start a run yourself, for instance after changing material, product data or your own rules. A run costs model calls, which is why starting it sits with the owner and admin roles.

on demand
Weekly Mondays at 04:00.

One fixed run per week, scheduled outside business hours. If something has shifted since the last run, you see it on Monday morning.

weekly
Negative feedback A thumbs down starts a recheck.

If somebody rates an answer in the chat negatively, exactly that question is sent through pipeline and judge panel again in the background. The question is already stored without personal details.

automatic
Traffic light

A run ends red, amber or green. Red as soon as at least one hard check failed, or the run itself broke off. Amber when there were only soft warnings. Otherwise green. In the detail view the red cases sit at the top, and open red cases without a decision count as a task in the Action Center.

What the Test Center does not do

So you rely on the right thing.

Your own rules

For a test question of your own you can note what the advisor should do, and optionally pick an expected reaction class. That expectation is not matched automatically today. The question runs and gets scored, the expectation is there for display.

History

The report shows the latest run. There is no history across several runs and no trend curve, and we do not pretend otherwise.

False alarms

A case you consider scored too strictly currently cannot be marked as a false alarm and muted for future runs. The data model provides for that state, the path to it is still missing.

Do not mix these up The ongoing comparison of published answers against your material is Safe Guard, not the Test Center. The Test Center asks the advisor questions, Safe Guard watches existing answers.
Next step

Put the advisor through the checks first.

One run, one traffic light, one list with the red cases on top. After that you know what to talk about instead of guessing it.