Refactor
First, uni-agent will focus on refactoring at the framework level to make the framework more user-friendly.
Then, building upon the new feat, we plan to implement the following features.
Performance optimization
These performance optimization features for Agentic RL will be initially incubated and validated within the uni-agent repository for rapid iteration. Once the architecture matures and stabilizes, we plan to upstream/migrate them into the main verl repository.
- Communication Bottleneck Optimization for Multi-Modal Agentic RL
- Hybrid Sequence Parallelism for Long-Context Agentic RL
Gateway
Agent
- Claude Code Harness
- Support multi-agent/multi-policy training
- Claw-style agent, e.g., openclaw
- AI4S agents, e.g., claude science
Task
- Support GUI Agent (e.g., OS-World)
Sandbox
RL recipe
Refactor
First, uni-agent will focus on refactoring at the framework level to make the framework more user-friendly.
Then, building upon the new feat, we plan to implement the following features.
Performance optimization
These performance optimization features for Agentic RL will be initially incubated and validated within the uni-agent repository for rapid iteration. Once the architecture matures and stabilizes, we plan to upstream/migrate them into the main verl repository.
Gateway
Agent
Task
Sandbox
RL recipe