Open Source LLM toolkit to build trustworthy LLM applications. TigerArmor (AI safety), TigerRAG (embedding, RAG), TigerTune (fine-tuning)
-
Updated
Dec 2, 2023 - Jupyter Notebook
Open Source LLM toolkit to build trustworthy LLM applications. TigerArmor (AI safety), TigerRAG (embedding, RAG), TigerTune (fine-tuning)
[NeurIPS 2024 Oral] Aligner: Efficient Alignment by Learning to Correct
[CoRL'23] Adversarial Training for Safe End-to-End Driving
AI-Generated Video Detection via Perceptual Straightening (NeurIPS2025)
Materials for the course Principles of AI: LLMs at UPenn (Stat 9911, Spring 2025). LLM architectures, training paradigms (pre- and post-training, alignment), test-time computation, reasoning, safety and robustness (jailbreaking, oversight, uncertainty), representations, interpretability (circuits), etc.
[ACL 2025 Findings] Fraud-R1 : A Multi-Round Benchmark for Assessing the Robustness of LLM Against Augmented Fraud and Phishing Inducements
Website to track people, organizations, and products (tools, websites, etc.) in AI safety
The go-to API for detecting and preventing prompt injection attacks.
Can Large Language Models Solve Security Challenges? We test LLMs' ability to interact and break out of shell environments using the OverTheWire wargames environment, showing the models' surprising ability to do action-oriented cyberexploits in shell environments
A high-performance string formatter written in Rust. This project detects and blocks LLM prompt injection and jailbreak attacks. It also features a customizable rule-based system and defends against obfuscated prompt attacks.
An official repository for the Capability-Based Scaling Laws for LLM Red-Teaming paper.
[NeurIPS 2024] SACPO (Stepwise Alignment for Constrained Policy Optimization)
A benchmark for evaluating hallucinations in large visual language models
Multi-agent simulation using LLMs. Agents autonomously decide actions for survival, reproduction, and social behavior in a grid world.This project aims to replicate a paper published in 2025 (arXiv:2508.12920).
Learned Semantic Decoder for Language Models.- Its the little model that sits under a big model's hat to explain what its thinking, just like the little cat's from Cat in the Hat! VOOM > FOOM
The official github repo for [LinguaSafe paper](https://arxiv.org/abs/2508.12733)
To associate your repository with the aisafety topic, visit your repo's landing page and select "manage topics."