Verifiers

Library for building verifiable RL environments and evals that run against any OpenAI-compatible endpoint.

Open SourceSelf HostedOffline Capable
0.0 (0)

About

Will Brown's verifiers grew from a GRPO experiment into a standard library for defining RL environments that LLMs train against, now maintained by Prime Intellect. An environment bundles a dataset, a rollout harness, and a scoring rubric into an installable Python module, and the same object serves three jobs: an eval suite runnable against any OpenAI-compatible endpoint, an agent harness with multi-turn tool use, and a training environment consumed by trainers like prime-rl for large-scale GRPO runs. Environment types span single-turn, multi-turn, and tool-calling agents, with sandboxed execution available for code tasks. Hundreds of community environments published on Prime Intellect's Environments Hub install with one command, which has turned the format into a de facto interchange standard for RL tasks and agent evals. The library itself is MIT licensed, requires Python 3.11 or newer, installs with pip install verifiers, and needs no GPU: evaluation talks to remote or local inference servers, and heavy training is delegated to a separate trainer.

Should you use Verifiers?

Pick it when

Pick it when you need reusable RL environments or agent evals, single-turn, multi-turn, or tool-calling, that run against any OpenAI-compatible endpoint today and feed a GRPO trainer such as prime-rl later, with no local GPU.

Look elsewhere when

It does not train models itself, so you still need prime-rl or another trainer for the heavy lifting. If you only want standard benchmark scores, lm-evaluation-harness or Inspect are more established evaluation tools.

Alternatives to Verifiers

  • OpenPipe ART

    Includes the trainer, GRPO onto LoRA adapters served by vLLM, so you train agents end to end rather than only defining environments.

  • Inspect AI

    Mature eval framework with 200+ prebuilt evaluations and a log viewer, aimed at measurement rather than RL training.

  • LM Evaluation Harness

    Standard runner for 60+ academic benchmarks; fixed tasks rather than custom environments with scoring rubrics.

  • TRL

    Has its own GRPO trainer with reward functions in Python; less portable than verifiers environments, but training is built in.

Reviews (0)

Leave a Review

No reviews yet. Be the first to review!

Details

Price
Free
Platform
Local/Desktop
Difficulty
Intermediate (3/5)
License
MIT
Added
Aug 24, 2026

Tags