Co-authors
6
list available
Resume
Academic Achievements
- Developed ProPS and ProPS+ methodologies to generate parameterized RL policies directly from LLMs using linguistic and numerical reasoning, enhanced by closed-loop feedback for in-context learning; outperformed state-of-the-art RL methods across 15 tasks.
- Contributed to OpenThought fine-tuned models that match DeepSeek-R1 performance on benchmarks like AIME and LiveCodeBench.
- Designed Evalchemy, a one-stop evaluation platform supporting over 30 benchmarks for downstream tasks including coding, mathematical reasoning, and instruction following.
- Co-authored paper 'Prompted Policy Search: Reinforcement Learning through Linguistic and Numerical Reasoning in LLMs' (submitted to NeurIPS 2025 Main Track).
- Co-authored 'OpenThoughts: Data Recipes for Reasoning Models' (arXiv:2506.04178, 2025).