SudhaKavya Bodapati Venkata

Talk Title:

When AI Becomes the Operator: Building Trustworthy Self-Healing Systems

Abstract:

Artificial intelligence is moving beyond monitoring systems and providing recommendations. Emerging AI-enabled systems can now analyze production incidents, identify possible causes, suggest corrective actions, roll back failed deployments, restore services, and automate selected operational tasks.

This talk explores how artificial intelligence, cloud automation, CI/CD pipelines, observability, and reliability engineering can work together to create self-healing technology systems. It will explain how these systems can detect problems, respond to failures, reduce downtime, and improve operational reliability in healthcare, manufacturing, and enterprise environments.

The talk will also address an important question: when should AI only assist engineers, and when can it be trusted to act independently?

Allowing AI to make operational decisions can improve speed and efficiency, but it can also introduce risks, including incorrect actions, cascading failures, excessive permissions, limited explainability, and decisions based on incomplete information. The session will present a practical approach for building trustworthy self-healing systems using human approval, confidence thresholds, audit trails, controlled permissions, limited impact boundaries, automated rollback, and continuous verification. Through practical examples, the talk will show how organizations can use AI responsibly to build systems that are more reliable, resilient, secure, and capable of recovering from failures with reduced manual effort.

Profile:

Sudhakavya Bodapati Venkata is a DevOps Engineer at Starkey Hearing Technologies and a Senior Member of IEEE. She specializes in building secure, scalable, and reliable cloud platforms for regulated healthcare and medical device manufacturing environments. Her work spans Microsoft Azure, Terraform-based infrastructure automation, CI/CD modernization, platform engineering, cloud security, and enterprise software delivery.

Her research focuses on using artificial intelligence to improve the reliability and intelligence of modern technology systems. Her work has explored predictive infrastructure orchestration, autonomous recovery of failed build and deployment pipelines, and safety-controlled approaches for using large language models in incident response. Her research has been presented at IEEE and Springer conferences. She also serves as a reviewer for international journals and conferences, contributing to the evaluation of research in artificial intelligence, cloud computing, DevOps, cybersecurity, and intelligent systems. She holds a master’s degree in Information Technology and Management from The University of Texas at Dallas. By combining hands-on engineering experience with applied research, her work advances the development of technology systems that are more resilient, secure, intelligent, and dependable.