Prompt Optimization Changes Which LLM is Best @DaveBennettTech
Prompt Optimization Changes Which LLM is Best  @DaveBennettTech
Uploaded May 2026 | Updated September 2026, 3 weeks ago
I tested Gemma 3 4B vs Ministral 8B on an intent classification task with the same prompt. Gemma 3 4B won. Then I optimized the prompt for Ministral using DSPy's MIPROv2 optimizer and ran both models again. Ministral pulled ahead. Gemma got worse.

A new paper from SAP and Stanford ("Optimization before Evaluation") found this happens across the board.

Code for my experiment: github.com/DaveBben/prompt-optimization-experiment
Paper: arxiv.org/abs/2604.27637
Dataset: bitext/Bitext-customer-support-llm-chatbot-training-dataset

— My Setup —
Models: Gemma 3 4B (Q8), Ministral 8B (Q8)
Inference: Llama.cpp on M4 Mac Pro
Optimizer: DSPy MIPROv2 with GPT-OSS 20B
Task: 27-class intent classification, 270 samples, stratified sampling

— Timestamps —
0:00 The benchmark problem
0:36 My experiment setup
1:14 First results — Gemma vs Ministral
1:25 Prompt optimization with DSPy
2:04 Round two — the ranking flip
2:46 What the paper found
3:45 What this means for model selection

#llm #dspy #promptoptimization #benchmarks #gemma #mistral #mlengineering #edgeai #machinelearning
Prompt Optimization Changes Which LLM is BestUse Android as MicrophoneBuilding LIDAR Smart Glasses with LasersGaming Emulation on ChromebookFirefox Quantum vs Chrome: The Real Results5 Actually Cool Android TricksHow to Play 4K YouTube videos on AndroidInstall Minecraft on Chromebook - 2021Back to School/College Apps & GadgetsHello, Galaxy S10Screen Mirror iOS to AndroidInstall Linux Apps on Chromebook
Dave Bennett |

Prompt Optimization Changes Which LLM is Best

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER