Uploaded August 2026 | Updated September 2026, 2 weeks ago
Guardrails are the seatbelt — the thing meant to hold if something goes wrong. Evaluation asks whether an agent gives good, accurate, grounded answers. Red teaming is a different question entirely: can you make the agent misbehave? This session shifts the mindset from "does this agent work" to "can I break it" — deliberately attacking your own agent before a real user, or a real attacker, does it for you.
We'll walk through the key categories of adversarial testing: jailbreaks, prompt injection, data exfiltration attempts, policy evasion, and role abuse.
We'll get concrete with real examples — like an instruction hidden in a PDF's white text, invisible to a human reviewer but not to the model processing it.
We'll also link back to guardrails: a red team failure isn't just a log entry, it points directly to the specific guardrail rule that should have caught it, turning every failure into an actual policy fix.
We’ll cover how Ejento's platform has been tested against more than 20,000 adversarial prompts, with automated red team runs executing on every deployment, not just once before launch, covering six OWASP LLM risk categories out of the box.
We will share our experience and take questions from the audience.
**What you'll learn**
- Why red teaming is a fundamentally different question from guardrails or evaluation — and why you need all three
- The key categories of adversarial testing: jailbreaks, prompt injection, data exfiltration, policy evasion, and role abuse
- How LLM red teaming differs from traditional penetration testing — a conversational attack surface, not a system vulnerability
- Why automated red team runs need to happen on every deployment, not as a one-time pre-launch report
- How a red team failure ties directly back to a specific guardrail rule, closing the loop between finding a gap and fixing it
- A live look at the red teaming dashboard — block rates, attack category breakdowns, and drill-down into a specific jailbreak — inside the Ejento platform
Guardrails are the seatbelt — the thing meant to hold if something goes wrong. Evaluation asks whether an agent gives good, accurate, grounded answers. Red teaming is a different question entirely: can you make the agent misbehave? This session shifts the mindset from "does this agent work" to "can I break it" — deliberately attacking your own agent before a real user, or a real attacker, does it for you.
We'll walk through the key categories of adversarial testing: jailbreaks, prompt injection, data exfiltration attempts, policy evasion, and role abuse.
We'll get concrete with real examples — like an instruction hidden in a PDF's white text, invisible to a human reviewer but not to the model processing it.
We'll also link back to guardrails: a red team failure isn't just a log entry, it points directly to the specific guardrail rule that should have caught it, turning every failure into an actual policy fix.
We’ll cover how Ejento's platform has been tested against more than 20,000 adversarial prompts, with automated red team runs executing on every deployment, not just once before launch, covering six OWASP LLM risk categories out of the box.
We will share our experience and take questions from the audience.
**What you'll learn**
- Why red teaming is a fundamentally different question from guardrails or evaluation — and why you need all three
- The key categories of adversarial testing: jailbreaks, prompt injection, data exfiltration, policy evasion, and role abuse
- How LLM red teaming differs from traditional penetration testing — a conversational attack surface, not a system vulnerability
- Why automated red team runs need to happen on every deployment, not as a one-time pre-launch report
- How a red team failure ties directly back to a specific guardrail rule, closing the loop between finding a gap and fixing it
- A live look at the red teaming dashboard — block rates, attack category breakdowns, and drill-down into a specific jailbreak — inside the Ejento platform









![Should You Trust ChatGPT With Your Data? | Jerry Liu x Data Science Dojo
🎙️ Future of Data and AI Podcast: Highlight with Jerry Liu (CEO & Co-Founder, LlamaIndex)
Should you trust ChatGPT with your data? Jerry Liu breaks it down.
In this highlight, Jerry explains how modern AI systems handle user data, what actually gets stored, and why understanding data flows is crucial before pasting sensitive information into any AI tool. He clarifies common misconceptions, privacy boundaries, and what organizations should keep in mind when using LLMs for real-world work.
💡 Key takeaway: AI tools aren’t inherently risky — but you need to know how they treat your data before you trust them.
Watch this clip to understand the real story behind data privacy in ChatGPT and other LLMs.
🔗 Watch the full episode: [Insert Link]
🎧 Explore more episodes: https://www.youtube.com/playlist?list=PL8eNk_zTBST_jMlmiokwBVfS_BqbAt0z2 Should You Trust ChatGPT With Your Data? | Jerry Liu x Data Science Dojo](https://i.ytimg.com/vi/nEDvHwM15mc/mqdefault.jpg)
