Hybrid AI Training Across Public and Private Data

Large-scale AI training relies on external compute, but private data may need to remain local. We explore a hybrid approach that trains on public data externally and private data locally, combining them while keeping private data and updates within the local environment.

Share this post

Choose a social network to share with, or copy the URL to share elsewhere

This is a representation of how your post may appear on social media. The actual post will vary between social networks

Preview

Abstract

Training a powerful AI model requires diverse training examples and substantial computing resources. Public datasets provide broad coverage, while private records capture cases specific to an institution and help address gaps in public data. When local computing capacity is insufficient, renting external servers provides additional resources, but conventional training requires transferring the training data to those servers. This creates a barrier
when private records must remain within the institution, preventing both sources from being brought together externally. We study a hybrid approach in which two models of the same architecture train in parallel: one learns from public data on external servers, while the other learns from private data locally. The private environment periodically imports the public model and averages its learned parameters with those of the private model,
combining their learning without sending private data or training updates outside. In an image-classification experiment, the hybrid approach raises private test accuracy from 22.08% with public-only training to 87.54%, while accuracy across the combined public and private test images increases from 80.03% to 81.88%. Further experiments show that the gains are strongest when private examples fill gaps in public coverage, and that the
averaging weights control the balance between private and overall performance. These results demonstrate that combining external public training with local private learning can substantially improve performance on locally important cases while preserving overall accuracy.