About the session
Why do smaller models perform well on simple training data, yet struggle when that data becomes more sophisticated? In this DevRev Live session, Ahmed Bashir, Ritesh Goru, and Shanay Mehta share lessons from fine-tuning language models for negotiation advice โ a nuanced task where quality cannot be checked with a simple right-or-wrong test.
The discussion follows the team's progression from a dataset generated with Gemini 2.5 Pro to more advanced training data produced using Gemini 3 Pro and Sonnet 4.5. They explore why changing base models and training settings did not close the performance gap, how training loss helps screen experiments, and the trade-offs between larger models, reasoning, and response latency.
Key takeaways
- Understand how stronger teacher models raise both dataset quality and performance expectations.
- Learn why rejection sampling and task-specific evaluators matter when preparing training data.
- Recognize when training loss signals that a smaller model is struggling to learn a more nuanced task.
- Explore the trade-offs between model size, reasoning, and inference latency.
- Distinguish improvements from more training data from persistent weaknesses in logic and phrasing.
Agenda
Distillation and initial fine-tuning results
The initial approach to training smaller models using frontier-model examples.
Stronger teacher models and more nuanced data
How improved teacher models raised the performance bar.
Comparing base models and training settings
What the team learned by changing models and training parameters.
Model capacity and training-loss signals
Using evaluators and training loss to diagnose learning limitations.
Larger models versus reasoning
The trade-offs between response quality and latency.
Dataset quality, size, and next experiments
Rejection sampling, data volume, and remaining weaknesses in logic and phrasing.
Speakers

Ahmed Bashir
Chief Technology Officer, DevRev
- RG
Ritesh Goru

Shanay Mehta
MTS @ DevRev




