Evals & Observability (14)
Evaluation, tracing and observability for agents and LLM applications.
Machine-readable: /api/evals-observability.json.
- AgentOps
Python SDK and platform for AI agent monitoring, LLM cost tracking, benchmarking, testing and debugging.
- Arize Phoenix
Open-source AI observability and evaluation platform from Arize.
- Braintrust
AI observability platform for agents: trace production, run evals and catch regressions before users see them.
- DeepEval
Open-source LLM evaluation framework with 50+ plug-and-play metrics for agents, RAG and chatbots.
- Helicone
Open-source AI gateway and LLM observability platform for routing, monitoring, evaluating and experimenting.
- Inspect
Open-source framework for large language model and agent evaluations from the UK AI Security Institute.
- Laminar
Open-source observability platform purpose-built for AI agents: trace, evaluate and debug agent failures.
- Langfuse
Open-source agent evals and observability platform: trace, evaluate and improve LLM applications and agents.
- LangSmith
LangChain's agent and LLM observability and evaluation platform: tracing, monitoring, cost and latency tracking.
- MCP Inspector
Official interactive developer tool for testing and debugging MCP servers in the browser, on the command line or in the terminal.
- OpenLLMetry
Traceloop's open-source observability for GenAI/LLM applications, built on OpenTelemetry.
- Opik
Comet's open-source platform to debug, evaluate and monitor LLM apps, RAG systems and agentic workflows.
- Promptfoo
Open-source tool to test and red-team prompts, agents and RAG: automated evals, vulnerability scanning and model comparison.
- W&B Weave
Weights & Biases toolkit for tracking, testing and improving LLM-powered applications.