Model Evaluation
Test the AI feature you are shipping, not a leaderboard model: tests written from your own description, run against your endpoint, graded with the reasoning shown and reported with confidence intervals.
One-time run
Has your chatbot changed? Run the same tests after a model or prompt update, or against every endpoint you run, and see exactly which answers regressed.
- Price
From $39per evaluation
- Starter, up to 300 tests$39
- Standard, up to 1,500 tests$129
- Scale, up to 3,000 tests$189
- You provide
- A description of what the assistant is for, and its endpoint
- Tests
- Up to 3,000, written from your description or imported as CSV or JSON
- You get back
- Pass rates with 95% confidence intervals, every failure with its reasoning, and PDF, Markdown or JUnit exports
- Subscription
- Not needed
Describe what your assistant does, must never do and should decline. Codity writes tests across ordinary, edge-case, ambiguous, out-of-scope and rule-breaking requests, in 15 categories from correctness to prompt injection and PII. Or bring up to 3,000 of your own as CSV or JSON.
OpenAI-compatible APIs, models in your own AWS Bedrock account, a tool on an MCP server, or any HTTP API. Credentials are encrypted, and requests only go to public HTTPS endpoints.
Every answer is judged against a written criterion, with the reasoning shown next to it. Critical tests, close calls and disagreements go to a three-judge panel, and a judge can abstain rather than guess.
Pass rates come with 95% confidence intervals, overall and per category. Tests whose repeated runs disagree are reported as flaky rather than counted as a pass or a fail. Export any run as PDF, Markdown or JUnit.
Run the suite again after a prompt or model change and Codity names every test that regressed, separates real change from noise with a paired statistical test, and fails the comparison on any critical or high-severity regression.
Tests · support-assistant
1,500 testsMore Features
- ReviewsContext-aware pull request review with summaries, requirement tracking, re-reviews and autofix.
- Security ScansSecrets, injection, auth gaps, dependency risk and an org-wide SBOM, caught before the merge.
- GovernanceMerge gates, policy checks and review rules, enforced on every pull request.
- Code NavigationA live architecture map, component health from Repo X-Ray, and answers across every repository.
- Developer AnalyticsReview latency, rework, throughput and DORA metrics per team and repository.
- Monitoring & InsightsContinuous repository monitoring with anomaly detection and deployment insight.
- Repo ScanA one-off, whole-repository review by eight specialist reviewers.
- PentestingBlack-box and active DAST testing of your live application, with evidence-checked findings.

