post-training

PackAI & models

Lets your agent train and align language models with human preferences using RLHF-style reinforcement learning.

Unavailable. Delivery for this kind is on the roadmap — not serving yet.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

About this app

RLHF and preference alignment including TRL, GRPO, OpenRLHF, SimPO, verl, slime, miles, and torchforge. Use when aligning models with human preferences, training reward models, or large-scale RL training.

Signals

GitHub stars
13k
Forks
931
Last commit
Jun 2026
Advanced
Item type
plugin
Key
orchestra-research-ai-research-skills-post-training
Source
github.com/orchestra-research/ai-research-skills