Towards Reliable Financial Agents: How a 4B Model Outsmarted a 235B Giant | Snorkel AI @aicouncilconf
Towards Reliable Financial Agents: How a 4B Model Outsmarted a 235B Giant | Snorkel AI  @aicouncilconf
Uploaded June 2026 | Updated September 2026, 2 weeks ago
[2026 - DAY 1 - WORKSHOP] Large generalist models have excellent reasoning but this does not necessarily imply specialized knowledge and tool calling capabilities. They can still hallucinate column names, ignore constraints, and generate SQL that returns nonsensical results. The problem isn’t intelligence—it’s reliability and specialization.

In this talk we’ll show how a 4B model was fine-tuned to outperform a 235B model on real financial analysis tasks. The key was not adding more reasoning ability, but enforcing tool discipline. Using synthetic data generation and reinforcement learning with the open-source rLLM framework, the model learned to explore schemas, validate outputs, and retry failures instead of hallucinating confident nonsense.

One key result: tool-use fundamentals generalize. Training on simple tool interactions transferred to much harder, multi-step financial tasks. If you’re building LLM systems that interact with databases, APIs, or internal tools, this talk focuses on the behaviors that actually matter — and how to teach them without frontier-scale compute.

SPEAKER:
Charles Dickens - Senior Research Scientist, Snorkel AI

👉 Sign up for our "No BS" Newsletter to get the latest technical data & AI content: aicouncil.com/newsletter

ABOUT AI COUNCIL:
AI Council brings together the brightest minds in data to share industry knowledge, technical architectures and best practices in building cutting edge data & AI systems and tools.

FIND US:
Website: aicouncil.com
LinkedIn: linkedin.com/company/aicouncilconf
X: https://x.com/aicouncilconf
Towards Reliable Financial Agents: How a 4B Model Outsmarted a 235B Giant | Snorkel AIFrom Spans to Trajectories: Observability for Long-Running Agents | HoneyHiveRLVR in Practice: From Synthetic Data to GRPO | NVIDIAOptimizing Model Training End-to-End: A Tiny MoE Case Study LambdaAI Launchpad 2026: Golden AnalyticsChang She on Why He Walked Away from Parquet to Build LanceDBThe 2% gains that 10x your inference | Lessons from AWS on optimizing VLMsQ&A with Scott Breitenother, Kilo: Engineers need to be the CEOs of agents. Are they ready?Beyond MLOps: Building AI systems with MetaflowAgentic AI: From Risk Awareness to Practical Control | Noma SecurityShould agents be durable? | RenderGuardrails for the Future AI Safety and Responsible AI in Practice
AI Council |

Towards Reliable Financial Agents: How a 4B Model Outsmarted a 235B Giant | Snorkel AI

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER