Uploaded June 2025 | Updated September 2026, 2 weeks ago
If the future is agentic, what does this mean for Infrastructure?
Moderated by Karthik Lakshminarayanan
- Delia David, an engineer at Meta who has been working on data infrastructure for recommendation systems and generative AI
- Anna Berenberg, an engineering fellow at Google Cloud who is responsible for the platform that supports Google's 300 products
- Barr Moses, the CEO and co-founder of Monte Carlo, a company that provides data and AI observability solutions
- Qi Ke, the corporate vice president at Microsoft Azure, responsible for Kubernetes and other cloud native services
Key Points
- Agentic systems, like AI agents, have fundamentally different characteristics compared to traditional software, requiring a rethinking of reliability, latency, and security in infrastructure.
- The rise of AI and agentic systems is disrupting traditional approaches to networking, load balancing, caching, and policy management, creating new challenges for infrastructure teams.
- Ensuring the reliability and security of AI-powered systems is critical, as failures can have significant business and brand implications.
- Observability and the ability to understand the behavior and provenance of agentic systems is essential for debugging and maintaining these systems.
- Humans will continue to play a crucial role in this new paradigm, with a need for "human in the loop" approaches to decision-making and action-taking.
Main Arguments
- Agentic systems, like AI agents, have fundamentally different characteristics compared to traditional software, requiring a rethinking of infrastructure (04:40, 06:35).
- Ensuring the reliability and security of AI-powered systems is critical, as failures can have significant business and brand implications (06:35, 12:51).
- Observability and the ability to understand the behavior and provenance of agentic systems is essential for debugging and maintaining these systems (09:47, 15:09).
Supporting Evidence
- Examples of failures, such as a chatbot selling a car for $1 and a chatbot hallucinating a non-existent policy (06:35).
- Challenges with networking, load balancing, caching, and policy management in the context of agentic systems (09:55).
- The need for comprehensive testing and evaluation of agentic systems to identify and mitigate potential failures (15:09, 19:51).
Learn more about the @Scale conference: atscaleconference.com/events/scale-data-ai-infra
If the future is agentic, what does this mean for Infrastructure?
Moderated by Karthik Lakshminarayanan
- Delia David, an engineer at Meta who has been working on data infrastructure for recommendation systems and generative AI
- Anna Berenberg, an engineering fellow at Google Cloud who is responsible for the platform that supports Google's 300 products
- Barr Moses, the CEO and co-founder of Monte Carlo, a company that provides data and AI observability solutions
- Qi Ke, the corporate vice president at Microsoft Azure, responsible for Kubernetes and other cloud native services
Key Points
- Agentic systems, like AI agents, have fundamentally different characteristics compared to traditional software, requiring a rethinking of reliability, latency, and security in infrastructure.
- The rise of AI and agentic systems is disrupting traditional approaches to networking, load balancing, caching, and policy management, creating new challenges for infrastructure teams.
- Ensuring the reliability and security of AI-powered systems is critical, as failures can have significant business and brand implications.
- Observability and the ability to understand the behavior and provenance of agentic systems is essential for debugging and maintaining these systems.
- Humans will continue to play a crucial role in this new paradigm, with a need for "human in the loop" approaches to decision-making and action-taking.
Main Arguments
- Agentic systems, like AI agents, have fundamentally different characteristics compared to traditional software, requiring a rethinking of infrastructure (04:40, 06:35).
- Ensuring the reliability and security of AI-powered systems is critical, as failures can have significant business and brand implications (06:35, 12:51).
- Observability and the ability to understand the behavior and provenance of agentic systems is essential for debugging and maintaining these systems (09:47, 15:09).
Supporting Evidence
- Examples of failures, such as a chatbot selling a car for $1 and a chatbot hallucinating a non-existent policy (06:35).
- Challenges with networking, load balancing, caching, and policy management in the context of agentic systems (09:55).
- The need for comprehensive testing and evaluation of agentic systems to identify and mitigate potential failures (15:09, 19:51).
Learn more about the @Scale conference: atscaleconference.com/events/scale-data-ai-infra










