What matters in AI.

Subscribe

Learn / AI basics

Definition · AI basics

Federated learning

Federated learning is a way of training one shared machine learning model across many devices or organisations without collecting their data in one place. Each participant, called a client, trains the model on its own data and sends only a model update to a central server, which combines the updates into the shared model.

Last reviewed

Key points

  • Federated learning trains one shared model without collecting the training data in one place. Each client trains locally and sends only a model update.
  • A central server picks clients, sends them the current model, and combines their updates into the next version.
  • Keeping raw data on the client reduces what is exposed, but the baseline setup gives no formal privacy guarantee. An update can still reveal something about the data behind it.
  • The server never receives a client's data, only its update, and a compromised client can send a malicious one. NIST names federated learning as one of the two settings where model poisoning attacks are most prevalent.

How it works

Google researchers named federated learning in a paper presented in 2017. Their target was phones: data too private or too large to upload, used to train a shared model anyway.

A central server runs training in rounds. In each round, as described by Kairouz et al.:

  1. The server picks a set of eligible clients.
  2. Each client downloads the current model.
  3. Each client trains it on its own data and computes an update.
  4. The server collects and combines the updates.
  5. The server changes the shared model using the combined update.

The raw training data stays on the client. The client sends back only its update.

Kairouz et al. describe two common settings, and say they are examples, not a complete list. Cross-device federated learning runs on many phones or other devices; Google uses it for its Gboard keyboard. Cross-silo federated learning joins a small number of relatively reliable clients, such as organisations that want a model trained on all their data but cannot share the data directly.

Why it matters

For privacy, it helps but does not guarantee. Keeping data on the client means less is collected. Kairouz et al. say the baseline setup still offers no formal privacy guarantee. Their example is a constructed scenario in which a server that knows the previous model and a client’s update would be able to infer one of the client’s training examples.

For security, a client can send a poisoned update. The server receives updates, not data, and a compromised client can send a malicious one to poison the shared model. NIST says model poisoning attacks are most prevalent in federated learning and in supply-chain attacks. NIST counts sending malicious updates as control of the model itself, a different capability from the training-data control that data poisoning needs.

NIST adds that aggregation rules which try to spot and exclude malicious updates exist, but motivated adversaries can bypass them.

Questions and answers

Is federated learning private?

More private than pooling the data, but not private by guarantee. Raw data stays on each client, yet Kairouz et al. say the baseline setup has no formal privacy guarantee. Their example is a constructed scenario in which knowing the previous model and a client's update would let the server infer one of that client's training examples.

Why is federated learning linked to model poisoning?

Because clients send the server model updates, and a compromised client can send a malicious one. NIST names federated learning, where "clients send local model updates to the aggregating server", as one of the two settings where model poisoning attacks are most prevalent. The other is the supply chain, where a supplier adds malicious code to a model.

Sources

  1. Communication-Efficient Learning of Deep Networks from Decentralized DataMcMahan, Moore, Ramage, Hampson and Agüera y Arcas (Google), AISTATS 2017, 17 Feb 2016
  2. Advances and Open Problems in Federated LearningKairouz, McMahan et al. (arXiv), 10 Dec 2019
  3. NIST AI 100-2e2025, Adversarial Machine Learning, A Taxonomy and Terminology of Attacks and MitigationsNational Institute of Standards and Technology, 24 Mar 2025