Evidence-constrained commercialization assessment with deterministic retrieval, six LLM stages, auditable scoring and checkpoint recovery.
-
Updated
Sep 14, 2026 - Python
Evidence-constrained commercialization assessment with deterministic retrieval, six LLM stages, auditable scoring and checkpoint recovery.
HuggingEnvs — RL Environments 101: building and scaling RL environments in the age of LLMs
Notus is a collection of fine-tuned LLMs using SFT, DPO, SFT+DPO, and/or any other RLHF techniques, while always keeping a data-first approach
Agentic RL 零基础中文教程:24 章从概念到 GRPO 实战,含 TRL 最小可跑示例 | Beginner-friendly Agentic RL tutorial with hands-on GRPO project
Distill teacher chains-of-thought into a LoRA adapter via a strict boxed-answer format contract + two-phase Train→Nudge (silver-medal NVIDIA Nemotron reasoning recipe, as a tested library).
An implementation of GRPO for Unsloth's VLMs training
Code repository dedicated to experimenting and research with tiny reasoning language model
Various training, inference and validation code and results related to Open LLM's that were pretrained (full or partially) on the Dutch language.
simpleR1: A Simple Framework for Training R1-like Models
An open-source application that estimates an Open Source Software Technology Readiness Level (OSSTRL) 1–9 from evidence that can be gathered automatically from a GitHub repository.
Training-data memorization auditor for fine-tuned LLMs — Trainer/TRL plugin, canary MIA + regurgitation audit, Apache-2.0
This project demonstrates the process of fine-tuning the Qwen2.5-3B-Instruct model using GRPO (Generalized Reward Policy Optimization) on the GSM8K dataset.
使用trl、peft、transformers等库,实现对huggingface上模型的微调。
Different post-training techniques for LLMs, including: SFT, DPO and Online RL
Vibe Innovation und Vibe Coding Workshop mit GitHub Codespaces
Build task models that replace frontier API calls
To associate your repository with the trl topic, visit your repo's landing page and select "manage topics."