post-training
PackAI & modelsLets your agent train and align language models with human preferences using RLHF-style reinforcement learning.
Unavailable. Delivery for this kind is on the roadmap — not serving yet.
Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
About this app
RLHF and preference alignment including TRL, GRPO, OpenRLHF, SimPO, verl, slime, miles, and torchforge. Use when aligning models with human preferences, training reward models, or large-scale RL training.
Signals
- GitHub stars
- 13k
- Forks
- 931
- Last commit
- Jun 2026
Advanced
- Item type
- plugin
- Key
orchestra-research-ai-research-skills-post-training- Source
- github.com/orchestra-research/ai-research-skills