Hospitals can train an AI model together without pooling their patient data, but a compromised participant can send malicious updates that undermine the shared model. Warden detects and contains that threat while training continues.
Our prototype uses Flower and PyTorch to train a pneumonia X-ray classifier across three simulated hospitals using real public medical images. Two training runs start identically, letting viewers compare protected and unprotected models side by side.
During live training, we can activate an attack at one hospital. A numerical monitor flags unusually large updates, and a rule-based security agent inspects the evidence and excludes the suspicious contribution before aggregation.
The dashboard shows training images, model predictions, measured accuracy, and the agent’s actions as they happen. It demonstrates real containment of a deliberately obvious model-poisoning attack, with a clear view of both the damage and the protection.