llm-eval-simple is a simple LLM evaluation framework with intermediate actions and prompt pattern selection
-
Updated
Feb 28, 2026 - Python
llm-eval-simple is a simple LLM evaluation framework with intermediate actions and prompt pattern selection
Synthetic datasets, experiment protocols, and evaluation code for "Governed Memory: A Production Architecture for Multi-Agent Workflows"
Experiments and Analyses for FilBench: An Open LLM Leaderboard for Filipino (EMNLP Main '25)
Official repository for "CausalDS: Benchmarking Causal Reasoning in Data-Science Agents"
Failed-subscription recovery agent for Indian payment rails. Benchmarked against 6 strategies and a hidden oracle over 200 paired seeds; derives the exact regulatory penalty (₹888.67/violation) at which compliance starts paying.
ArxivRoll tells you “How much of your score is real, and how much is cheating?” AAAI'26 Code of paper: How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework
Governed autonomous research agent for evaluating fine-tuned reasoning critics — sandboxed execution, provenance tracking, and blinded adjudication.
Reproducible benchmark framework for testing hypotheses about AI coding agents
Evaluating a RAG pipeline with DeepEval — Answer Relevancy, Faithfulness, Contextual Precision, Recall & Relevancy scored by an LLM judge
An automated LLM evaluation and benchmarking framework using LangChain, DeepEval, and PostgreSQL. Features a real-time Streamlit dashboard to analyze model accuracy, latency, token usage, and hallucination rates.
Benchmark and evaluation harness for LLM-based parallel code translation across CUDA, OpenMP, OpenCL, and OpenMP target offload (NeurIPS 2026)
To associate your repository with the llm-evaluation-benchmark topic, visit your repo's landing page and select "manage topics."