Skip to content

Repository files navigation

Agent Flight Recorder

An open-source, OpenTelemetry-native reliability platform for AI agents. Capture production agent behavior, replay failures, turn traces into regression tests, evaluate changes, and maintain an auditable record of every agent decision and tool call.

Reliability Loop

trace → replay → eval → regression test → policy check → audit trail

Agent Flight Recorder is not primarily another LLM observability dashboard. It focuses on the workflow enterprises need before trusting agents with meaningful autonomy.

Quick Start (Target Experience)

Python:

from agent_flight_recorder import recorder

recorder.init(
    app_name="support-agent",
    environment="development"
)

with recorder.agent_run(name="refund-agent", user_id="user_123") as run:
    result = agent.invoke("Refund my latest order")

TypeScript:

import { recorder } from "@agent-flight-recorder/node";

recorder.init({
  appName: "support-agent",
  environment: "development",
});

await recorder.agentRun(
  {
    name: "refund-agent",
    userId: "user_123",
  },
  async () => {
    return await agent.invoke("Refund my latest order");
  }
);

See docs/quickstart.md for full setup.

Run locally

cp .env.example .env
make setup
make dev

See CONTRIBUTING.md for native (non-Docker) development.

Verify the full Phase 1 loop:

make e2e

Run CI regression tests (Phase 2):

make test
make policy-test
make storage-test   # requires Docker (Postgres + ClickHouse + MinIO)
make integration-test  # LangGraph + OpenAI Agents example traces
make prod-up        # production compose stack

CLI examples:

afr replay run_abc123 --model gpt-4.1-mini
afr eval run examples/evals/refund_tool_correctness.yml --run-id run_abc123
afr test ./examples/afr-tests/
afr policy check <run_id>
afr policy load examples/policies/require_approval_for_large_refunds.yml

Documentation

Doc Description
quickstart.md Install, configure, and capture your first agent run
architecture.md System components, data model, and tech stack
replay.md Replay modes, snapshots, and CLI usage
evals.md Evaluator types, YAML config, and CI regression tests
policies.md Policy rules, risk detection, and violation handling
integrations.md LangGraph and OpenAI Agents SDK tracing helpers

Architecture Decisions

ADR Title
ADR-001 Build an OpenTelemetry-Native Agent Flight Recorder
ADR-002 Storage Strategy for Agent Traces and Replay Data
ADR-003 Redaction and Privacy Model
ADR-004 Evaluation and Regression Testing Strategy
ADR-005 Open Source Core vs. Commercial Cloud Boundary

Full index: adr/README.md

MVP Scope

Must Have

Python & TypeScript SDKs, manual instrumentation, OTel-compatible traces, local collector, SQLite storage, trace timeline UI, model/tool call capture, cost/latency/error tracking, basic redaction, replay from stored trace, manual evals, trace-to-regression-test conversion, Docker Compose, demo support agent.

Should Have

OpenAI Agents SDK & LangGraph integrations, OTLP export, basic policy checks, prompt/model replay comparison, JSON/YAML eval config, GitHub Actions regression example.

Could Have

ClickHouse/Postgres storage, hosted cloud, team accounts, SSO/RBAC, Slack alerts, advanced PII detection, MCP governance, LLM-as-judge evals.

Implementation Phases

Phase Focus Exit Criteria
1 Local trace capture ✅ Capture agent run locally; inspect model/tool calls; cost/latency/redaction/search/replay/eval; make e2e passes
2 Replay & regression ✅ afr CLI; model replay; afr test CI gate; GitHub Actions regression workflow
3 Policy & risk ✅ Policy YAML; tool risk + PII detection; violation UI; make policy-test
4 Production storage & export ✅ Postgres + ClickHouse + MinIO; OTLP/Langfuse/Phoenix export; make storage-test
5 Integrations & dashboard ✅ LangGraph + OpenAI Agents helpers; /dashboard UI; make integration-test

Key Architectural Bet

OpenTelemetry should be the compatibility layer, but agent-specific replay, evaluation, policy, and audit semantics should be the differentiation layer.

Repository layout

apps/collector      FastAPI ingestion API
apps/web            Next.js trace viewer
packages/sdk-js     TypeScript SDK
packages/sdk-python Python SDK
packages/cli          afr CLI (replay, eval, test)
packages/shared-schema  Span types and JSON schemas
examples/           Demo agents
infra/              Docker Compose and SQLite schema

License

Apache License 2.0.

You may self-host, modify, and embed the SDKs in proprietary applications. A separate hosted or enterprise offering (SSO, RBAC, managed retention, support SLAs) may be offered commercially without restricting the open-source core. See ADR-005.

About

OpenTelemetry-native reliability platform for AI agents: trace, replay, eval, and audit.

Resources

Contributing

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages