tokens&
For enterprises
Submit
Sign in
tokens&

Build better AI stacks, claim useful opportunities, and give AI infrastructure companies a source-labeled adoption readout they can trust.

For buildersFor enterprises

Product

  • For builders
  • Category rankings
  • Startup credits and perks
  • Agent Skills
  • Platform
  • Submit project, tool, product, or perk

Enterprise

  • Start free company workspace

Community

  • Community
  • Newsletter
  • Events
Xin

© 2026 tokensand, LLC. All rights reserved.

  • Terms
  • Privacy
  • Security
  • Data Processing
  • Status
  1. Home
  2. Tools
  3. Observability

Observability Rankings

Ranked by AgentRank — adoption proof, trust evidence, and protocol readiness

Unclaimed LeaderAgentRank: 36
Langfuse verified adoptionsTrust 44

This tool leads the category but has not been claimed by its maker.

Claim this tool to see who is evaluating it
#ToolAgentRankTrustVerified7d VelocityProtocols
Langfuse3644SuppressedSuppressedAPI, SDK
Helicone3644SuppressedSuppressedAPI, SDK
AnomalyArmor2246SuppressedSuppressedAPI, SDK
Phoenix2144SuppressedSuppressedAPI, SDK
LangWatch2144SuppressedSuppressedAPI, SDK
Laminar2144SuppressedSuppressedAPI, SDK
Opik2144SuppressedSuppressedAPI, SDK
Future AGI2144SuppressedSuppressedAPI, SDK
TensorZero2144SuppressedSuppressedAPI, SDK
TruLens2144SuppressedSuppressedAPI, SDK
Giskard2144SuppressedSuppressedAPI, SDK
Evidently2144SuppressedSuppressedAPI, SDK
Pydantic Logfire2144SuppressedSuppressedAPI, SDK
BoundFlow2144SuppressedSuppressedAPI, SDK
MLflow2144SuppressedSuppressedAPI, SDK
OpenInference2144SuppressedSuppressedAPI, SDK
Ragas2144SuppressedSuppressedAPI, SDK
AgentOps2144SuppressedSuppressedAPI, SDK
Inspect AI2144SuppressedSuppressedAPI, SDK
OpenLIT2144SuppressedSuppressedAPI, SDK
Langfuse logo
Langfuse

Observability· langfuse.com

Freemium

Open-source LLM engineering platform for tracing, evals, prompt management, and metrics so teams can debug and improve production AI applications.

Why recommended

Based on popular builder stacks until you save tools or publish a project.

29,121
Open source
MLflow logo
MLflow

Observability· mlflow.org

Open source

MLflow is an open-source AI engineering platform with GenAI evaluation, tracing, monitoring, and optimization workflows for agents, LLM apps, and models.

Why recommended

Based on popular builder stacks until you save tools or publish a project.

Helicone logo
Helicone

Observability· helicone.ai

Freemium

Open-source AI gateway and LLM observability platform for logging, tracing, routing, evaluation, caching, and cost monitoring across model providers and agent stacks.

Why recommended

Based on popular builder stacks until you save tools or publish a project.

AgentOps logo
AgentOps

Observability· agentops.ai

Freemium

Developer platform for tracing, debugging, and deploying AI agents with SDK-based observability, session replay, and framework integrations.

Why recommended

Based on popular builder stacks until you save tools or publish a project.

Promptfoo logo
Promptfoo

Observability· promptfoo.dev

Open source

Open-source testing, red-teaming, and evaluation framework for prompts, RAG, and agents.

Why recommended

Based on popular builder stacks until you save tools or publish a project.

Opik logo
Opik

Observability· comet.com

Open source

Open-source LLM observability and evaluation platform from Comet for tracing, prompt management, testing, and agent optimization.

Why recommended

Based on popular builder stacks until you save tools or publish a project.

Evidently logo
Evidently

Observability· evidentlyai.com

Freemium

Open-source observability and evaluation framework for monitoring ML and LLM applications with tests and metrics.

Why recommended

Based on popular builder stacks until you save tools or publish a project.

Guardrails AI logo
Guardrails AI

Observability· guardrailsai.com

Freemium

LLM validation and guardrail framework for checking structured outputs, policy compliance, and production AI failure modes before release.

Why recommended

Based on popular builder stacks until you save tools or publish a project.

AnomalyArmor logo
AnomalyArmor

Observability· anomalyarmor.ai

Freemium

Data observability platform with hosted and local MCP access for alerts, freshness, schema drift, lineage, and quality checks.

Why recommended

Based on popular builder stacks until you save tools or publish a project.

Giskard logo
Giskard

Observability· giskard.ai

Freemium

Open-source evaluation and testing library for LLM agents, RAG systems, and AI risk checks.

Why recommended

Based on popular builder stacks until you save tools or publish a project.

AgentAcct logo
AgentAcct

Observability· github.com

Open source

Local-first work-intelligence dashboard for coding agents that attributes tokens, estimated cost, tasks, and verified evidence from local session logs.

Why recommended

Based on popular builder stacks until you save tools or publish a project.

BoundFlow logo
BoundFlow

Observability· boundflow.github.io

Open source

Open-source control plane for governed AI agent execution with lifecycle policies, approvals, durable runs, audit receipts, and observability.

Why recommended

Based on popular builder stacks until you save tools or publish a project.

Kiln logo
Kiln

Observability· kiln.tech

Open source

Open-source AI engineering workbench for model evaluation, dataset curation, fine-tuning, and app reliability workflows.

Why recommended

Based on popular builder stacks until you save tools or publish a project.

OpenPipe ART logo
OpenPipe ART

Observability· art.openpipe.ai

Open source

Agent Reinforcement Trainer for improving LLM agent reliability with GRPO-based reinforcement learning workflows.

Why recommended

Based on popular builder stacks until you save tools or publish a project.

Future AGI logo
Future AGI

Observability· futureagi.com

Freemium

Open-source agent engineering platform that combines evaluations, observability, experiments, and guardrails so teams can ship and improve production AI agents faster.

Why recommended

Based on popular builder stacks until you save tools or publish a project.

Ragas logo
Ragas

Observability· ragas.io

Open source

Evaluation framework for testing RAG and LLM applications with metrics, synthetic test data, and CI-friendly quality checks.

Why recommended

Based on popular builder stacks until you save tools or publish a project.

Portkey logo
Portkey

Observability· portkey.ai

Freemium

AI gateway and observability layer for routing, guardrails, caching, logging, and spend control across LLM providers behind one API.

Why recommended

Based on popular builder stacks until you save tools or publish a project.

Inspect AI logo
Inspect AI

Observability· inspect.aisi.org.uk

Open source

Framework for large language model evaluations, task definitions, solvers, scoring, and reproducible eval runs.

Why recommended

Based on popular builder stacks until you save tools or publish a project.

Phoenix logo
Phoenix

Observability· arize.com

Open source

Open source AI observability and evaluation platform for tracing LLM applications, running evals, and debugging agent behavior with OpenTelemetry.

Why recommended

Based on popular builder stacks until you save tools or publish a project.

Braintrust logo
Braintrust

Observability· braintrustdata.com

Freemium

AI observability and evals platform for tracing production systems, running experiments, and catching regressions.

Why recommended

Based on popular builder stacks until you save tools or publish a project.

Laminar logo
Laminar

Observability· laminar.sh

Freemium

Open-source observability, tracing, and evaluation platform built for AI agents.

Why recommended

Based on popular builder stacks until you save tools or publish a project.

LangWatch logo
LangWatch

Observability· langwatch.ai

Freemium

The complete LLMOps platform for agent testing, evaluations, prompt management, and production observability across AI applications.

Why recommended

Based on popular builder stacks until you save tools or publish a project.

TensorZero logo
TensorZero

Observability· tensorzero.com

Open source

Open-source LLM application stack for observability, evaluations, prompt optimization, and controlled model gateway workflows.

Why recommended

Based on popular builder stacks until you save tools or publish a project.

MCPJam Inspector logo
MCPJam Inspector

Observability· mcpjam.com

Freemium

Testing and debugging workbench for MCP servers, MCP apps, and ChatGPT apps.

Why recommended

Based on popular builder stacks until you save tools or publish a project.

OpenLIT logo
OpenLIT

Observability· openlit.io

Freemium

OpenTelemetry-native AI engineering platform for LLM observability, evaluations, guardrails, and prompt operations.

Why recommended

Based on popular builder stacks until you save tools or publish a project.

OpenInference logo
OpenInference

Observability· arize-ai.github.io

Open source

OpenTelemetry semantic conventions and instrumentation packages for tracing LLM calls, agent steps, tool use, and retrieval workloads.

Why recommended

Based on popular builder stacks until you save tools or publish a project.

DeepEval logo
DeepEval

Observability· deepeval.com

Open source

Open-source evaluation framework for testing LLM apps, agents, RAG systems, and prompts with repeatable metrics, regression checks, and CI-friendly workflows.

Why recommended

Based on popular builder stacks until you save tools or publish a project.

Traceloop OpenLLMetry logo
Traceloop OpenLLMetry

Observability· traceloop.com

Open source

OpenTelemetry-based open-source instrumentation for tracing, monitoring, and debugging LLM applications across frameworks and providers.

Why recommended

Based on popular builder stacks until you save tools or publish a project.

Pydantic Logfire logo
Pydantic Logfire

Observability· logfire-us.pydantic.dev

Freemium

Observability platform for production Python, LLM, and agent systems from the Pydantic team.

Why recommended

Based on popular builder stacks until you save tools or publish a project.

TruLens logo
TruLens

Observability· trulens.org

Open source

Open-source evaluation and tracking toolkit for measuring LLM app quality, feedback functions, and RAG behavior.

Why recommended

Based on popular builder stacks until you save tools or publish a project.

Auralogs logo
Auralogs

Observability· auralogs.ai

Freemium

Agent-focused observability platform that exposes structured production logs and AI analyses to Codex, Claude, Cursor, and other clients through read-only MCP.

Why recommended

Based on popular builder stacks until you save tools or publish a project.

Open source
Alluvia logo
Alluvia

Observability· github.com

Open source

Local-first CLI that mines Claude Code, Cursor, and ChatGPT histories into searchable engineering insights without uploading conversations.

Why recommended

Based on popular builder stacks until you save tools or publish a project.

Scorable by Root Signals logo
Scorable by Root Signals

Observability· scorable.ai

Freemium

LLM and agent evaluation platform with reusable scorers, test datasets, and SDKs for measuring production AI behavior.

Why recommended

Based on popular builder stacks until you save tools or publish a project.

Open source
Snyk Agent Scan logo
Snyk Agent Scan

Observability· github.com

Open source

Open-source Snyk scanner for AI agents, MCP servers, and agent skills that detects prompt injection, tool poisoning, toxic flows, and agent supply-chain risks.

Why recommended

Based on popular builder stacks until you save tools or publish a project.

mcpsnoop logo
mcpsnoop

Observability· github.com

Open source

Transparent MCP proxy and terminal inspector for observing tool calls, payloads, latency, and failures between clients and servers.

Why recommended

Based on popular builder stacks until you save tools or publish a project.

Weights & Biases Weave logo
Weights & Biases Weave

Observability· wandb.ai

Open source

LLM observability and evaluation toolkit for tracing, debugging, and improving production agents and AI applications.

Why recommended

Based on popular builder stacks until you save tools or publish a project.

Narrow the observability decision

Head-to-head comparisons and adoption evidence for the products ranked above.

Compare the observability shortlist

  • Langfuse vs MLflowSide-by-side pricing, API surface, and adoption evidence.
  • Langfuse vs HeliconeSide-by-side pricing, API surface, and adoption evidence.
  • MLflow vs HeliconeSide-by-side pricing, API surface, and adoption evidence.
  • Langfuse alternativesWhat developers evaluate in place of Langfuse.
  • MLflow alternativesWhat developers evaluate in place of MLflow.
  • Helicone alternativesWhat developers evaluate in place of Helicone.

Check the adoption evidence

  • Observability developers actually keep usingFiltered to products with verified usage events, cohort breadth, and retention.
  • AgentRank across every categoryHow each category is ordered by adoption proof, trust, and protocol support.
  • Trending developer toolsProducts gaining verified adoption fastest right now.
  • Best AI Observability ToolsShortlist guide for this category, with the tradeoffs written out.
26,852
Open source
5,808
Open source
5,639
Open source
22,625
Open source
19,688
Open source
7,650
Open source
7,081
Open source
1
Open source
5,476
Open source
424
Open source
5
Open source
4,937
Open source
10,353
Open source
1,192
Open source
14,555
Open source
12,067
Open source
2,296
Open source
10,176
Open source
21
API available
3,039
Open source
3,307
Open source
11,673
Open source
2,034
Open source
2,562
Open source
1,050
Open source
16,317
Open source
7,240
Open source
4,350
Open source
3,400
Open source
2
Open source
2,764
Open source
267
Open source
1,102
Open source