Skip to content
#

ai-evaluation-framework

Here are 35 public repositories matching this topic...

prompt-evaluator is an open-source toolkit for evaluating, testing, and comparing LLM prompts. It provides a GUI-driven workflow for running prompt tests, tracking token usage, visualizing results, and ensuring reliability across models like OpenAI, Claude, and Gemini.

  • Updated Dec 4, 2025
  • TypeScript

Frontend for Calibrate, a framework for evaluating AI agents: speech-to-text, text-to-speech, LLM evaluation, end-to-end simulations

  • Updated Sep 14, 2026
  • TypeScript

👽 An Alien Mind — The Epistemic Operating System for the AI Age. Open-source epistemic layer between humans and AI: Cognitive Firewall, multidimensional Trust Profile, five-agent Alien Council, Thought DNA provenance, behavioral fingerprints, and Monte Carlo simulation. Not a chatbot. Not a guardrail. Not a lie detector.

  • Updated Sep 8, 2026
  • Python

Add this topic to your repo

To associate your repository with the ai-evaluation-framework topic, visit your repo's landing page and select "manage topics."

Learn more