THE AFTER-CONFERENCE PROCEEDING OF THE AIC 2026 WILL BE SUBMITTED FOR INCLUSION TO IEEE XPLORE

Keval Barvaliya

Keval Barvaliya

Mercurial Cores and Corrupted Gradients: Silent Data Corruption as an Unmeasured Failure Mode in Applied AI

Abstract:

Every guarantee an organization makes about an intelligent system whether it is reproducibility, auditability, verified agent behaviour, model provenance, all rests on an assumption that is rarely stated aloud: that the processor underneath returned the correct answer. Fleet-scale reliability studies from hyperscale operators have shown this assumption to be unsafe, documenting processors and accelerators that produce incorrect results without crashing, without raising a machine check, and often only under specific inputs, instruction paths, or thermal and voltage conditions. Machine learning is an unusually poor detector of such faults, because it has no oracle: there is no ground truth against which a corrupted gradient, a wrong reduction, or a mis-quantized weight can be checked, so the fault is absorbed as a marginally worse model rather than surfacing as an error. This keynote traces silent data corruption through the applied AI stack - training and gradient aggregation, checkpoint integrity, quantization and compilation, and multi-tenant inference serving and then examines numerical non-determinism, where identical weights and identical inputs diverge as a function of batch composition and kernel scheduling. It treats the adversarial case through published Rowhammer-class bit-flip attacks that degrade trained networks toward random-guess accuracy by corrupting a small number of bits in memory, positioning hardware fault injection as an attack surface that never touches the model interface. Drawing a connection to oracle-free verification in probabilistic inference, where convergence must be established diagnostically rather than confirmed against a known answer, the talk closes on practical mitigation - fleet screening and core qualification, redundant and checksum-based execution, batch-invariant kernels, and corruption-aware checkpointing and asks what auditable AI can honestly mean when the substrate itself is probabilistic.

Profile:

Keval Barvaliya is a Senior Integration Consultant at Tyson Foods and an independent researcher working at the intersection of systems engineering and advanced computation. In his professional role he designs cloud-native systems, multi-layer integration patterns, and high-reliability distributed architectures at enterprise scale. His published research centres on reliability and interpretability in probabilistic and large-scale AI systems, including rank-aware, tail-sensitive convergence diagnostics for multi-chain inference and spectral-entropy measures of representational complexity in transformer residual streams, alongside work on reversible computing foundations in quantum computation and cloud-tiered pipelines for built-environment data, published through IEEE ETCOM, Atlantis Press, and the International Journal of Computer Applications. He received a Best Paper Award at the 7th International Conference on Data Engineering and Communication Technology for Verified Contract Execution through Automated Analysis. He is a Member of IEEE and the IEEE Computer Society, a Full Member of Sigma Xi, an ACM Certified Peer Reviewer, and an editorial board member of the International Journal of Computer Science Trends and Technology, and has served as a judge for the ACM NextGen Hackathon 2026 and the New Jersey School Boards Association STEAM TANK Challenge. This keynote extends the question running through his convergence-diagnostics work - how to establish that a computation can be trusted when no ground-truth oracle exists - downward into the compute substrate itself: silent data corruption, numerical non-determinism, and hardware-level fault injection. 

© Copyright @ aic2026. All Rights Reserved