Your AI system
The production endpoint or agent ASHE evaluates under real query and task patterns.
AI Reliability Infrastructure
Evaluate RAG systems and AI agents with evidence-backed testing that exposes failures before your users do.
The problem
A polished response does not prove retrieval worked, an agent completed its task, or the answer stayed grounded in evidence. Teams need verdicts they can inspect, not chat logs they have to trust.
Looks correct
Polished answer, wrong underlying retrieval
Unsupported claim
Confident response with no source support
Incomplete retrieval
Relevant documents never surfaced
Hallucinated detail
Specific fact not present in knowledge base
Wrong source
Answer cites irrelevant retrieved document
Inconsistent
Same question, different answers across runs
Why basic testing fails
Basic pass/fail testing doesn't explain why failures happen. ASHE evaluates correctness, retrieval, hallucination resistance, and case-level diagnosis.
01
A test can pass.
Hand-picked questions often look fine in isolation.
02
But retrieval can still fail.
The right documents may never surface under real query patterns.
03
The answer can sound convincing.
Fluent language hides unsupported claims and missing evidence.
04
The system can hallucinate.
Specific facts appear that were never in the knowledge base.
AI reliability infrastructure
One pipeline for RAG reliability and agent reliability: scenario, execution, evidence, judgment, and verdict.
Your AI system
The production endpoint or agent ASHE evaluates under real query and task patterns.
Query or task
ASHE generates scenarios from your knowledge base or task set and sends them to the target.
Context capture
Retrieved documents, tool calls, and execution traces are recorded when the system returns them.
Response
The system output is judged for correctness, grounding, and task completion.
Evaluation
Frontier intelligence scores quality dimensions and flags cases that need deeper review.
Evidence
Case-level evidence chains expose retrieved sources, traces, and judge reasoning.
Verdict
Structured pass/fail outcomes with deterministic failure classification.
Introducing ASHE
ASHE connects to your knowledge base or agent task sets, runs evaluations against live RAG and agent systems, and returns structured reports with evidence you can act on.
01
Your Knowledge Base
02
Generate evaluation questions
03
Execute against your system
04
Evaluate responses
05
Analyze evidence
06
Classify failures
07
Generate recommendations
08
Actionable evaluation report
What ASHE measures
RAG scoring weights correctness, retrieval, and hallucination resistance. Agent scoring adds task completion, tool use, and trajectory quality, with case-level evidence behind each dimension.
answer_correctness
45%
Correctness
Answer accuracy against expected knowledge from generated questions.
retrieval_relevance
30%
Retrieval
Relevance of returned documents when retrieval is available.
hallucination_resistance
25%
Hallucination resistance
Groundedness of the response against retrieved evidence.
Evaluation lifecycle
Scroll to traverse the full evaluation pipeline. Each stage stays pinned until the next locks in.
ASHE reads the supplied knowledge base and builds the evaluation context.
01 / 08Evaluation scenarios are generated from your knowledge base, not hand-picked demos.
02 / 08Scenarios are sent to your production system under realistic query and task patterns.
03 / 08Frontier intelligence evaluates answer and retrieval quality with structured verdicts.
04 / 08Low-confidence or ambiguous cases escalate to Advanced Reasoning for deeper review.
05 / 08Failures are classified deterministically instead of buried in aggregate scores.
06 / 08Evidence chains show why each verdict was reached, case by case.
07 / 08Actionable recommendations surface the next engineering step.
08 / 08Product workflow
Scroll to walk from project setup through evaluation, cases, evidence, and report.
01
Project
Create a project to organize targets and knowledge bases.
02
Target
Register the endpoint or agent ASHE will evaluate.
03
Knowledge base or task set
Upload ground-truth content for RAG questions, or define agent tasks with optional tools and policy.
04
Evaluation run
Start an async run. ASHE returns 202 and executes in the background.
05
Cases
Inspect individual failures with verdict and failure classification.
06
Evidence
See retrieved documents, judge reasoning, and ground-truth comparison.
07
Verdict
Structured outcome with deterministic failure categories.
08
Report
Overall quality, component breakdown, and trends across runs.
project.config
name: customer-support-ai
targets: 2
knowledge_bases: 1
target.endpoint
type: assistant | agent
url: /api/query
auth: bearer
evaluation.config
rag: refund-policy.md
agent: support_tasks.json
status: ready
evaluation.run
status: running
cases: 48 / 128
concurrency: 4
case
verdict: hallucination
failure: unsupported_claim
confidence: low
evidence_chain
retrieved: refund-policy.md
score: 0.91
ground_truth: match
verdict
result: failed
category: hallucination
review: required
evaluation_report
overall_quality: 82
correctness: 88
retrieval: 79
01
Project
Create a project to organize targets and knowledge bases.
project.config
name: customer-support-ai
targets: 2
knowledge_bases: 1
02
Target
Register the endpoint or agent ASHE will evaluate.
target.endpoint
type: assistant | agent
url: /api/query
auth: bearer
03
Knowledge base or task set
Upload ground-truth content for RAG questions, or define agent tasks with optional tools and policy.
evaluation.config
rag: refund-policy.md
agent: support_tasks.json
status: ready
04
Evaluation run
Start an async run. ASHE returns 202 and executes in the background.
evaluation.run
status: running
cases: 48 / 128
concurrency: 4
05
Cases
Inspect individual failures with verdict and failure classification.
case
verdict: hallucination
failure: unsupported_claim
confidence: low
06
Evidence
See retrieved documents, judge reasoning, and ground-truth comparison.
evidence_chain
retrieved: refund-policy.md
score: 0.91
ground_truth: match
07
Verdict
Structured outcome with deterministic failure categories.
verdict
result: failed
category: hallucination
review: required
08
Report
Overall quality, component breakdown, and trends across runs.
evaluation_report
overall_quality: 82
correctness: 88
retrieval: 79
Product preview
evaluation_report
82
Overall quality
Correctness
88
Retrieval
79
Hallucination
76
case_diagnosis
What is the refund policy for enterprise customers?
system_answer
Enterprise customers receive a 60-day refund window.
verdict: hallucination
failure: unsupported claim in retrieved evidence
recommendation: verify enterprise policy in knowledge base
Evidence
Every case exposes retrieved documents, judge reasoning, ground-truth comparison, and the evidence chain that led to the verdict.
Question → system response → evidence and trace
→ Ground truth → Judge assessment → Verdict
“Customers may request a refund within 30 days of purchase.”
source: refund-policy.md · score: 0.91
Failure diagnosis
FAILED
↓ why?
↓ retrieval_failure
↓ what evidence?
↓ what should I investigate?
Continuous evaluation
Who ASHE is for
AI Engineers
Validate system quality before every release.
ML Engineers
Investigate evaluation failures with structured evidence.
Product & Engineering Teams
Track reliability as the product evolves.
Teams shipping AI assistants
Know whether production behavior can be trusted.
Built for teams shipping AI
ASHE is early-stage infrastructure focused on honest evaluation — no inflated claims, no fake social proof. Real capabilities you can verify in the product.
RAG + Agent evaluation
Structured reliability testing for retrieval systems and single AI agents.
Evidence-backed scoring
Scores, verdicts, and supporting evidence — not opaque pass/fail labels.
OpenAI-powered evaluation
Frontier Intelligence, Advanced Reasoning, and Premium Reasoning routed by ASHE.
Production-oriented diagnostics
Failure categories, step-level analysis, and trace inspection for agents.
Evaluation infrastructure
Inside the product
One active workspace area at a time, from projects through reviews, synchronized with the preview panel as you scroll.
Projects
Organize targets and knowledge bases.
Evaluation Runs
Start and monitor evaluations.
Reports
Understand overall quality and trends.
Cases
Inspect individual failures.
Evidence
See why a case received its verdict.
Reviews
Mark cases requiring human attention.
About ASHE
01
The problem we care about
Production AI quality cannot be inferred from demo conversations. A polished answer does not prove retrieval worked, a task was completed, or that claims are supported.
02
What we believe
AI systems should be evaluated by evidence, not vibes. Systematic evaluation with transparent scoring is how teams ship with confidence.
03
Why evaluation needs more than pass/fail
Aggregate green checks do not expose failure categories, evidence chains, execution traces, or the reasons a case failed. Teams need structured insight, not a single score.
04
What ASHE is designed to do
ASHE generates evaluation scenarios, executes them against your system, judges outcomes with frontier model capabilities, and produces reports with weighted scoring and case-level evidence.
05
Where we're going
Repeatable AI reliability infrastructure for engineers and ML teams: track quality across iterations, diagnose failures, and improve with evidence.
Connect a RAG system or agent, define scenarios from a knowledge base or task set, and see where reliability actually stands.