Trace Claude Code sessions in Braintrust via the Braintrust Claude Plugin
-
Updated
Sep 2, 2026 - Python
Trace Claude Code sessions in Braintrust via the Braintrust Claude Plugin
1st Place Winner (General Judge) - Datadog Self-Improving Agents Hack. Two identical AI agents play Split or Steal. No pre-programmed betrayal. They discover deception on their own. Built with @evancorrea.
The braintrust pi extension now lives in a new monorepo: https://github.com/braintrustdata/braintrust-coding-agent-plugins
My implementation for a kaggle competition: https://www.kaggle.com/competitions/WattBot2025
Trace Codex sessions in Braintrust via the Braintrust Codex Plugin
Can you eval an art form? Canon is a continuity linter for serialized TV, YouTube and micro-drama fiction. Canon plays the role of whats currently the scriptwriting coordinator, verifies your story's logic and cites the scene. No GenAI writing here. That's left to the humans.
Braintrust Coding Agent Plugins
Trace Antigravity sessions in Braintrust via the Braintrust Antigravity Plugin
Orchestrate other AI CLIs (agy, Codex, Grok, OpenCode, Claude Code) for second opinions, research, and codebase analysis. No Gemini CLI. Hybrid always-on skill (eval-backed).
CacheCatch audits AI agent context and shows what to move so repeated tokens hit cache instead of full-price input — it finds context waste, prompt cache misses, and hidden agent cost leaks; and give you actionable insights how to fix them.
Dual-modular platform combating SRE alert burnout and securing generative AI deployments using Spring Boot and FastAPI. Tracks on-call fairness via Gini coefficients and uses AST/CST analysis to detect fragile, low-quality, or overly AI-dependent code in CI/CD pipelines while introducing quantifiable engineering contribution metric
Decision-grade comparison: LangSmith vs Langfuse vs Arize Phoenix vs Braintrust vs MLflow 3 for multi-team agent observability and evaluation. Versions, pricing and self-hosting terms verified against primary sources (August 2026).
LLM evaluation platform — MMLU knowledge benchmarks across 57 subjects plus agentic evaluation with real finance tools, scored by an LLM judge.
Trace Grok sessions in Braintrust via the Braintrust Grok Plugin
Learn LLM evals by fixing an app that's quietly lying to its customers. A dog daycare, a naive prompt, and nine steps with Braintrust.
AI-powered coloring page generator for kids, parents, and teachers.
AI agent that turns regulated enterprise conversations into decision-ready briefs and audited actions. Permission-aware RAG with grounded citations, deterministic policy gates, HITL approval loop, and a 3-vertical eval scorecard.
A causal debugger for distributed AI systems—isolated Daytona worlds, Braintrust evals, and evidence-first incident response.
A Python implementation of the canonical agent architecture: a while loop with tools. Build production ready AI agents with purpose-built tools, comprehensive tracing, and async patterns.
Bilingual LLM-as-a-judge benchmark and dashboard for measuring agreement, robustness, coverage, and when AI evaluators should abstain.
To associate your repository with the braintrust topic, visit your repo's landing page and select "manage topics."