Rubric-based RL that normalizes judged quality over only the responses satisfying the hard constraints.
🛵 Should they come looking for me, I inten
Aman Behera
beingamanforever
AI & ML interests
Long Horizon Agentic RL, OPD, Generative Engine Optimization, Performance Optimisation
Recent Activity
liked a model about 16 hours ago
numind/NuExtract3 liked a dataset 8 days ago
krutrim-ai-labs/ocr_rotation_bench liked a model 8 days ago
qualcomm/MobileNet-v3-SmallOrganizations
None yet