-
Natural Language Reinforcement Learning
Paper • 2411.14251 • Published • 30 -
Benjamin-eecs/Llama-3.1-8B-Instruct-NLRL-TicTacToe-Value
Feature Extraction • 8B • Updated • 15 -
Benjamin-eecs/Llama-3.1-8B-Instruct-NLRL-TicTacToe-Policy
Feature Extraction • 8B • Updated • 15 -
Waterhorse/Llama-3.1-8B-Instruct-NLRL-Breakthrough-Value
Feature Extraction • 8B • Updated • 17
🤝 Open to Collab
Bo Liu
AI & ML interests
None yet
Recent Activity
upvoted a paper about 11 hours ago
UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement published a dataset 7 days ago
Benjamin-eecs/openrsi-commit-runtime-assets updated a dataset 7 days ago
Benjamin-eecs/openrsi-commit-runtime-assets