Making bad CLIs fun with Small Language Models [PyCon DE & PyData 2026] @PyDataTV
Making bad CLIs fun with Small Language Models [PyCon DE & PyData 2026]  @PyDataTV
Uploaded August 2026 | Updated September 2026, 2 weeks ago
🔊 Recorded at PyCon DE & PyData 2026, 16.04.2026
2026.pycon.de/talks/GMNE3E

🎓 Watch Senior Data Scientist Moritz Bauer demonstrate how to transform frustrating command-line interfaces into intuitive, natural-language experiences using fine-tuned Small Language Models.

Speakers:
Moritz Bauer

Description:
Complex command-line interfaces (CLIs) often suffer from steep learning curves due to intricate syntax and obscure flags. While large language models (LLMs) can translate natural language into CLI commands, they typically require cloud connectivity, API keys, and significant computational resources. To address this, a local implementation using small language models (SLMs) with fewer than one billion parameters enables fast, on-device inference without internet dependency.

The approach centers on supervised fine-tuning (SFT) using synthetic datasets. Initial data pairs of natural language instructions and corresponding CLI arguments are generated by prompting a coding LLM (such as Claude Opus) with the tool's source code. To prevent the SLM from overfitting to rigid patterns, these pairs undergo prompt mutation, where a teacher model (such as Qwen 3.5 9B or GLM 4.5) generates multiple natural language variations for each command. To maintain data integrity, a filtering step removes variants that omit critical identifiers, such as customer names or dates, which would otherwise force the SLM to hallucinate.

The implementation utilizes the Gemma 2B or smaller variants (specifically a quarter-billion parameter model) fine-tuned via the Hugging Face stack on a MacBook M3 Pro. Training 100,000 pairs takes approximately 10 hours. Testing on an internal plotting tool and FFmpeg yielded an accuracy rate of roughly 85%, with errors typically manifesting as missing or extra flags. Using the Lama CPP inference engine, the system generates commands in under 2.5 seconds, demonstrating that SLMs can effectively map domain-specific natural language to complex technical syntax locally.

⭐️ About PyCon DE:
PyCon DE is the leading conference on open-source Python applications in AI and data science. It brings together industry professionals, researchers, AI and data science practitioners, and software engineering communities, providing a unique platform for collaboration, knowledge sharing, and innovation.

The PyCon DE & PyData 2026 conference delivered an exceptional experience, fostering stronger connections within the Python community while showcasing the latest advancements in artificial intelligence and data science. Attendees enjoyed a diverse and engaging program of talks, workshops, and networking opportunities, further establishing the conference as a premier event for Python, AI, and data science enthusiasts across Germany.

PyCon DE 2027 will take place in Heidelberg from 19 to 23 April 2027.

Follow us:
• Newsletter: 2027.pycon.de/newsletter
• LinkedIn: linkedin.com/company/pyconde
• X: x.com/pyconde

Links:
• Conference website: pycon.de
• Other sessions: 2026.pycon.de/talks

The conference was organized by
• Python Softwareverband e.V.: pysv.org
• Pioneers Hub gemeinnützige GmbH: pioneershub.org
in collaboration with NumFOCUS Inc.: numfocus.org


If you enjoyed this session, please like, and subscribe to our channel for more insightful talks and discussions.
Share this video with your network to spread the knowledge!

Hashtags:
#Python #PyConDE #PyData #OpenSource #AI #DataScience #MachineLearning #SoftwareEngineering #LLMs #Community #Sovereignty

Acknowledgements:
Special thanks to all the volunteers and sponsors who made this event possible.

About:
Python Softwareverband e.V.:
PySV is a non-profit that promotes the use and development of Python in Germany through events, education, and advocacy, fostering an open Python community.

Pioneers Hub gemeinnĂźtzige GmbH:
is a non-profit fostering innovation in AI and tech by connecting experts and promoting knowledge exchange through events and collaborative initiatives.

NumFOCUS Inc.
supports open-source scientific computing by providing financial and logistical support to key projects like NumPy and Jupyter, promoting sustainable development and collaboration.


pydata.org

PyData is an educational program of NumFOCUS, a 501(c)3 non-profit organization in the United States. PyData provides a forum for the international community of users and developers of data analysis tools to share ideas and learn from each other. The global PyData network promotes discussion of best practices, new approaches, and emerging technologies for data management, processing, analytics, and visualization. PyData communities approach data science using many languages, including (but not limited to) Python, Julia, and R.
Making bad CLIs fun with Small Language Models [PyCon DE & PyData 2026]AI is nowhere near as good as Twitter wants you to think.Doclings Viral 37,000 GitHub StarsNeal Richardson - MCP, or not MCP | Pydata London 26DeepSeek & The Emergent Aha! MomentUsing Sensor Fusion and ML to Navigate Underground When GPS Fails [PyCon DE & PyData 2026]The CEO of OSS4AI: dont stay on the surface.Why Co-Design and Transparency are Critical in Machine LearningTracking Knowledge Diversity in LLM-Generated Responses. [PyCon DE & PyData 2026]Avoiding Common Agent FailuresFoundation Models in Forecasting: Are We There Yet? Lessons from the TrenchesThe Day the Agent Started Lying (Politely) [PyCon DE & PyData 2026]
PyData |

Making bad CLIs fun with Small Language Models [PyCon DE & PyData 2026]

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER