Closing the Learning Loop: Human-in-the-Loop Active Learning for Self-Improving Enterprise AI Pipelines
Abstract:
Enterprise document extraction pipelines built on large language models often accumulate extensive human correction data without using those corrections to improve subsequent batches. This creates a structural learning gap in which recurring extraction errors persist while valuable signals remain isolated in QA systems.
This session presents a closed-loop Human-in-the-Loop (HITL) architecture that turns production corrections into continuous improvement signals. The approach unifies extraction events, human corrections, and prompt configuration changes in an append-only event store, then applies active learning strategies, primarily uncertainty and disagreement sampling, to focus human review on the instances most likely to generate informative corrections. Validated signals are propagated through a prompt evolution policy with A/B validation and automated regression protection.
The implementation uses Amazon Augmented AI (A2I) for review workflow management and Amazon SageMaker Ground Truth for annotation and auto-labeling, with AWS S3, Glue, Athena, Lambda, Step Functions, and SageMaker Pipelines supporting the feedback workflow. Dynamic confidence-threshold calibration and quality-gated auto-labeling further reduce unnecessary review.
The framework was evaluated across production enterprise document-extraction deployments, covering 130+ field types and six document categories over eight consecutive production batches. It reduced false-positive review traffic by 38% and increased the self-healing rate from 41% to 81%. Human review volume was reduced by approximately 55% during the evaluation period.
Attendees will learn practical patterns for designing unified correction data, selecting informative review instances, implementing validation-gated prompt evolution, managing cold-start and reviewer-quality challenges, and converting reactive human review into a measurable improvement mechanism for enterprise AI pipelines.
Profile:
Avneet Bansal is a Delivery Consultant at Amazon Web Services (AWS) Professional Services with 16+ years of experience in cloud infrastructure, DevOps engineering, and AI/ML solution delivery. He specializes in designing autonomous, self-healing AI agents for regulated, enterprise production systems, including a commercialized Intelligent Contract Processing Platform built on Amazon Bedrock AgentCore. He also contributes to open source through the textract-field-memory library (https://github.com/aws-samples/sample-textract-field-memory), which provides automated field learning and spatial anomaly detection for document processing pipelines. He holds 16 industry certifications, including 8 AWS certifications and CNCF's CKA, CKAD, and CKS, and has been recognized with AWS ProServe Delivery Excellence and Top Performer honors, along with an "Exceeds High Bar" performance rating every year since joining AWS. He has authored multiple peer-reviewed publications spanning generative AI, human-in-the-loop learning systems, distributed systems, and cloud security.
You can send your queries to the following email ID:
aic@scrs.in
+91-7503322444
(whatsapp messages only)
© Copyright @ aic2026. All Rights Reserved