Red teaming AI agents before they ship @Datasciencedojo
Red teaming AI agents before they ship  @Datasciencedojo
Uploaded August 2026 | Updated September 2026, 2 weeks ago
Guardrails are the seatbelt — the thing meant to hold if something goes wrong. Evaluation asks whether an agent gives good, accurate, grounded answers. Red teaming is a different question entirely: can you make the agent misbehave? This session shifts the mindset from "does this agent work" to "can I break it" — deliberately attacking your own agent before a real user, or a real attacker, does it for you.

We'll walk through the key categories of adversarial testing: jailbreaks, prompt injection, data exfiltration attempts, policy evasion, and role abuse.

We'll get concrete with real examples — like an instruction hidden in a PDF's white text, invisible to a human reviewer but not to the model processing it.

We'll also link back to guardrails: a red team failure isn't just a log entry, it points directly to the specific guardrail rule that should have caught it, turning every failure into an actual policy fix.

We’ll cover how Ejento's platform has been tested against more than 20,000 adversarial prompts, with automated red team runs executing on every deployment, not just once before launch, covering six OWASP LLM risk categories out of the box.

We will share our experience and take questions from the audience.

**What you'll learn**

- Why red teaming is a fundamentally different question from guardrails or evaluation — and why you need all three
- The key categories of adversarial testing: jailbreaks, prompt injection, data exfiltration, policy evasion, and role abuse
- How LLM red teaming differs from traditional penetration testing — a conversational attack surface, not a system vulnerability
- Why automated red team runs need to happen on every deployment, not as a one-time pre-launch report
- How a red team failure ties directly back to a specific guardrail rule, closing the loop between finding a gap and fixing it
- A live look at the red teaming dashboard — block rates, attack category breakdowns, and drill-down into a specific jailbreak — inside the Ejento platform
Red teaming AI agents before they shipRunning your LLM agent safely: Hands-on with Docker sandboxesFear of Change The Real AI Barrier #AIAdoption #ArtificialIntelligence #BiggestBarrierToAIThe ONE Trait That Separates Successful People From Everyone Else | Docker x Data Science DojoWhat Are Enterprises Using MCPs For? | João Moura x Data Science DojoTutorial: Building Event-Driven Agents with Docker | Future of Data and AI | Agentic AI ConferenceWorkshop: Building Flexible RAG Systems: RAG Retriever Optimization with LanceDBAgentic Behavior & Multi-Agent Basics | Multi Agent Workflows for Beginners | Part 1Whats The Biggest Mistake Enterprises Make When Adopting AIBuilding Trustworthy AI with NVIDIA NeMo Guardrails Live Demo #ai #guardrailsShould You Trust ChatGPT With Your Data? | Jerry Liu x Data Science DojoWorkshop: Building Smarter Agents, Faster with Arize | Future of Data and AI | Agentic AI Conference
Data Science Dojo |

Red teaming AI agents before they ship

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER