Uploaded November 2024 | Updated September 2026, 1 hour ago
JSON mode has been one of the biggest enablers for working with Large Language Models! JSON mode is even expanding into Multimodal Foundation models! But how exactly is JSON mode achieved?
There are generally 3 paths to JSON mode: (1) constrained generation (such as Outlines), (2) begging the model for a JSON response in the prompt, and (3) A two stage process of generate-then-format.
I am BEYOND EXCITED to publish the 108th Weaviate Podcast with Zhi Rui Tam, the lead author of Let Me Speak Freely? A Study on the Impact of Format Restrictions on Performance of Large Language Models!
As the title of the paper suggests, although constrained generation is awesome because of its reliability, we may be sacrificing the performance of the LLM by producing our JSON with this method.
The podcast dives into how these experiments identify this and all sorts of details about the potential and implementation details of Structured Outputs. I particularly love the conversation topic of incredible Complex Structured Outputs, such as generating 10 values in a single inference.
I hope you enjoy the podcast! As always please reach out if you would like to discuss any of these ideas further!
Links
Let me speak freely? - arxiv.org/abs/2408.02442
Let me speak freely (repo) - github.com/appier-research/structure-gen
StructuredRAG - arxiv.org/abs/2408.11061
StructuredRAG (repo) - github.com/weaviate/structured-rag
Are more LM Calls All You Need? - arxiv.org/pdf/2403.02419
Outlines - github.com/dottxt-ai/outlines
No more bad outputs with structured generation: Remi Louf - youtube.com/watch?v=aNmfvN6S_n4
Weaviate Gorilla - github.com/weaviate/weaviate-gorilla
Weaviate Podcast #88 with Jason Liu (Creator of Instructor) - youtube.com/watch?v=higlHgYDc5E
DSPy Typed Predictors from Thomas Ahle - github.com/stanfordnlp/dspy/blob/main/examples/functional/functional.ipynb
Chapters
0:00 Welcome Ray!
0:40 Motivation for Structured Outputs
2:35 Let Me Speak Freely?
5:10 JSON Prompting
8:03 Inference Scaling Perspectives
12:45 Weaviate Agents with Function Calling
13:50 Learning to output JSON
18:25 Experimental Analysis in Let Me Speak Freely?
24:18 Really Complex Structured Outputs
27:15 Reasoning with Structure
29:05 Directions for Benchmarking Structured Outputs
31:45 Function Calling
34:02 Importance of Ordering Outputs
36:30 Exciting Future Directions for AI
JSON mode has been one of the biggest enablers for working with Large Language Models! JSON mode is even expanding into Multimodal Foundation models! But how exactly is JSON mode achieved?
There are generally 3 paths to JSON mode: (1) constrained generation (such as Outlines), (2) begging the model for a JSON response in the prompt, and (3) A two stage process of generate-then-format.
I am BEYOND EXCITED to publish the 108th Weaviate Podcast with Zhi Rui Tam, the lead author of Let Me Speak Freely? A Study on the Impact of Format Restrictions on Performance of Large Language Models!
As the title of the paper suggests, although constrained generation is awesome because of its reliability, we may be sacrificing the performance of the LLM by producing our JSON with this method.
The podcast dives into how these experiments identify this and all sorts of details about the potential and implementation details of Structured Outputs. I particularly love the conversation topic of incredible Complex Structured Outputs, such as generating 10 values in a single inference.
I hope you enjoy the podcast! As always please reach out if you would like to discuss any of these ideas further!
Links
Let me speak freely? - arxiv.org/abs/2408.02442
Let me speak freely (repo) - github.com/appier-research/structure-gen
StructuredRAG - arxiv.org/abs/2408.11061
StructuredRAG (repo) - github.com/weaviate/structured-rag
Are more LM Calls All You Need? - arxiv.org/pdf/2403.02419
Outlines - github.com/dottxt-ai/outlines
No more bad outputs with structured generation: Remi Louf - youtube.com/watch?v=aNmfvN6S_n4
Weaviate Gorilla - github.com/weaviate/weaviate-gorilla
Weaviate Podcast #88 with Jason Liu (Creator of Instructor) - youtube.com/watch?v=higlHgYDc5E
DSPy Typed Predictors from Thomas Ahle - github.com/stanfordnlp/dspy/blob/main/examples/functional/functional.ipynb
Chapters
0:00 Welcome Ray!
0:40 Motivation for Structured Outputs
2:35 Let Me Speak Freely?
5:10 JSON Prompting
8:03 Inference Scaling Perspectives
12:45 Weaviate Agents with Function Calling
13:50 Learning to output JSON
18:25 Experimental Analysis in Let Me Speak Freely?
24:18 Really Complex Structured Outputs
27:15 Reasoning with Structure
29:05 Directions for Benchmarking Structured Outputs
31:45 Function Calling
34:02 Importance of Ordering Outputs
36:30 Exciting Future Directions for AI










- Star us on GitHub https://github.com/weaviate/weaviate
- Stay updated and subscribe to our newsletter: https://newsletter.weaviate.io/
- Try out Weaviate Cloud Services for free here: https://console.weaviate.cloud/
Got a question?
- Forum: https://forum.weaviate.io/
- Slack: https://weaviate.io/slack
Connect with us on
- Twitter: https://twitter.com/weaviate_io
- LinkedIn: https://www.linkedin.com/company/weaviate-io/ Generic Search is Dead.](https://i.ytimg.com/vi/Yn0ftG5Dh0c/mqdefault.jpg)
