AI Awesome
Home
Discover the Best AI Tools
Your ultimate directory for finding the right artificial intelligence solutions for any task.
Search
AI Testing
(71)
L
open source
LLM Debugger
LLM Debugger is a VS Code research extension that runs a program, sets breakpoints and inspects runtime values, using Jev for action choices and a text model for hypotheses and patches. It requires model credentials and actually executes target code. It is an unsupported proof of concept; small example benchmarks do not guarantee general debugging success.
AI Testing
B
open source
Browser Use Olympics
Provides browser task benchmarks and a bounded Jev automation loop over Chrome CDP. Requires browser and API setup; benchmark timing excludes planning and does not establish general performance.
AI Testing
J
open source
JevArena
Runnable spectator demo compares two Jev agents playing Snake through indexed browser controls. Real model play requires a TypeSafe key; without one it clearly uses deterministic rules. This narrow demonstration is not a general agent benchmark.
AI Testing
J
open source
Jev Frontend QA
Jev Frontend QA runs contract-based or exploratory browser checks with Jev, Browser Harness and optional host-planned subgoals. It correlates DOM, HTTP and persistence evidence into reports; exploration is not a contract pass. It needs Chrome and live TypeSafe access, while native selects, embedded contexts and popup workflows remain unsupported.
AI Testing
C
open source
CodexQA Jev Browser (openqa-cn)
CodexQA Jev Browser indexes page controls for goal-driven actions, test-case generation and model-free replay, retaining screenshots and visible assertions. Live goal runs need Jev or a chat API; it runs its own Playwright Chromium. Experimental site knowledge and action coverage limit success, and offline fixture tests do not prove live reliability.
AI Testing
R
unknown
Reticle
Reticle instruments an application so coding agents can verify flows using runtime state, network and console evidence and replay saved checks; app integration is required, cloud history sync is optional, and server/enterprise packages have different licensing restrictions from the Apache-licensed SDK.
AI Testing
A
open source
Agent Handoff Gate
An experimental protocol repository includes runnable offline tools for replaying metrics on a fixed historical 40-case evaluation and checking protocol fixtures. It supplies no runtime agent controller or live API evaluation, and the original model responses are unavailable; the replay is not a general-purpose evaluator.
AI Testing
D
open source
decision-first
decision-first combines an agent skill with Python runners for testing typed Jev decisions, comparing labels, and documenting reusable experiments. It requires Python and a TypeSafe key for live tests, sends supplied state to the API, and relies on human evaluation before adopting an integration.
AI Testing
T
unknown
typesafe-ai-firewall
typesafe-ai-firewall is a runnable shadow-evaluation harness for AI tool-call risk decisions, with Jev hazard scores, deterministic policy and offline log replay. It requires a TypeSafe key for live scoring and does not enforce production tool execution. Synthetic test success does not establish real-world protection; calibration and latency targets remain unmet.
AI Testing
J
open source
Jev Belay
Jev Belay is a Claude stop hook that checks local transcripts for edited files lacking later successful checks, then asks Jev whether completion should be blocked. It requires Jev credentials, skips turns without edits and lets execution finish on errors; it is not a comprehensive verifier.
AI Testing
Z
open source
Zevals
Zevals is a TypeScript library for testing complete agent conversations with assertions on responses and tool use, optional user simulation and model-based judges. It works with different test runners; model-driven evaluation needs provider credentials and does not guarantee correctness.
AI Testing
J
open source
jev-test-filter
Uses a Git diff and Jev judgments to select potentially affected tests and emit runner-specific filters, with optional execution and snapshot review. Requires Node.js, Git and a service key; code context is sent remotely and selective testing cannot replace all regression coverage.
AI Testing
P
open source
ProxyKit MCP
ProxyKit MCP lets agents inspect captured HTTP traffic and, with enabled permissions, control mocks, capture and replay through a local ProxyKit engine. The binary and engine are proprietary; this repository provides MIT-licensed documentation and distribution metadata.
AI Testing
Q
open source
qa-probe
qa-probe maps frontend API calls to backend routes and probes live endpoints to produce evidence-based diagnostic reports for developers and agents. Authenticated checks require authorized backend access; inconclusive results remain unknown rather than passing.
AI Testing
F
open source
Feather Wand
Feather Wand adds model chat, test-plan context, script refactoring and agent-driven edits to Apache JMeter, with optional Jev tool routing. It needs JMeter and configured model or CLI access. Suggestions and real test-plan edits require review; cost displays may be estimates, and anonymous usage telemetry is enabled by default with an opt-out.
AI Testing
J
open source
JsonFabrica MCP Server
JsonFabrica MCP lets coding agents manage templates and generate synthetic JSON records and related batches through the JsonFabrica API. The server runs locally over stdio, but generation uses the configured gateway, API credentials and service billing.
AI Testing
W
open source
wcagc-mcp
wcagc-mcp connects assistants to wcagc accessibility scans, PDF structure checks and findings through a thin API adapter. It requires service authentication and applicable plan access; automated results cover only part of accessibility and do not establish compliance.
AI Testing
Q
unknown
QAI Consultant
QAI Consultant creates QA strategies, risk registers, effort estimates and test plans, and offers deterministic document and results assessments. The full generation application needs configured model services and Pinecone; its local MCP tools can supply QA knowledge and calculations without those services.
AI Testing
A
open source
Ariadne
Ariadne uses sazanami’s Kotlin static analysis and Git changes to select affected tests for coding agents. Selection favors fast iteration and may miss indirect dependencies; it does not replace a full or conservative CI test run.
AI Testing
S
open source
Supercov
Supercov combines local test-coverage reporting with Jev-based quality and security judgments to guide coding-agent checks. Coverage uses the project’s test command; AI scoring requires a TypeSafe key, sends source for evaluation and is not a security certification.
AI Testing
I
open source
imagedimensions-mcp
imagedimensions-mcp audits natural versus rendered image dimensions, oversized assets and formats on public web pages. Scanning runs on imagedimensions.com rather than a local browser; it needs a reachable public URL and does not itself optimize the images.
AI Testing
I
open source
inkcheck
inkcheck mechanically compiles and explores ink stories to report runtime failures, reproduction paths and unreachable content. It needs Node.js 18+ and uses no AI itself; bounded exploration can be incomplete and cannot prove every story path is correct.
AI Testing
T
open source
Tarn
Tarn runs YAML-defined API tests and returns structured failures, with CLI and MCP tools for validation, execution and fix planning. Tests make real requests to configured endpoints. Shell steps require explicit opt-in; otherwise they are skipped and marked passed, so a successful run does not mean every declared step executed.
AI Testing
V
open source
VerifyAX MCP Server
VerifyAX MCP connects AI clients to the VerifyAX platform for registering agents, generating scenarios, running evaluations and reading results. A VerifyAX API key is required, and scenario generation and evaluations consume platform credits.
AI Testing
Previous
Page 1 of 3
Next