← Back to Blog

MLS Protocol: How Group Messaging Stays Forward-Secure

MLS Protocol: How Group Messaging Stays Forward-Secure - QNSQY post-quantum encryption guide

For more than a decade, secure messaging in groups was an open problem. The Signal protocol gave us excellent end-to-end encryption for two-person chats, with forward secrecy and post-compromise security baked in. Once the conversation involved ten people, then a hundred, then a thousand, the design fell apart. Naively encrypting separately to each member of a thousand-person group is wasteful and slow. The sender-key approaches like Signal's group sessions, WhatsApp's group ratchets, and Matrix's Megolm scale better but lose post-compromise security: a compromised member key keeps revealing future messages until the group rekeys.

MLS, the Messaging Layer Security protocol, was specified by the IETF MLS Working Group to solve this. After six years of design, MLS was published as RFC 9420 in July 2023. The architecture combines a continuous group key agreement with a logarithmic-cost ratcheting tree. A 10,000-person group can rekey in under a minute on a modern phone. A compromised member's keys leak nothing about the group state once the group rotates through its next epoch.

This article walks through what MLS is, why it matters, how the TreeKEM data structure works, and where post-quantum cryptography fits.

Why Groups Need Their Own Protocol

The Signal protocol's pairwise design is brilliant but does not group well. To send a message to a group of N people you would run N parallel Signal sessions, derive N separate session keys, and encrypt the message N times. That is fine for small groups; it falls over for large ones.

The "sender keys" trick used by WhatsApp, Matrix Megolm, and others maintains a single ratcheting key per sender. Each member receives the sender's chain key once and ratchets it locally as messages arrive. Encryption is just one AEAD per message. This scales well but has a weakness: when a member is added or removed, every other member's chain has to be redistributed (one Signal session per member to deliver the new key), and the cost of that distribution is O(N).

Worse, post-compromise security in sender-key designs is delayed by however long the group goes between key rotations. If the rotation interval is a week, a compromised member's keys keep working for up to a week. For high-stakes deployments — corporate boards, military command groups, hospital incident response teams — that is too long.

MLS solves both problems simultaneously. Group operations (add, remove, update keys) cost O(log N) work per participant. Post-compromise security advances on every group operation, not just on schedule. The group can churn aggressively without breaking down.

TreeKEM: The Core Idea

The technical heart of MLS is a data structure called TreeKEM, originally proposed by Bhargavan, Beurdouche, and others in 2018. TreeKEM is a binary tree where each leaf represents a member and each internal node represents a shared secret derivable by all members in that subtree.

When Alice wants to update her key, she generates a new keypair and a fresh path of secrets from her leaf up to the root. She encrypts the new path secrets to the public keys of the siblings along her copath (the nodes adjacent to her path). Every other member learns the secrets in their subtree and updates their tree state, but only Alice and the path nodes change. The cost of this update is O(log N) keypair operations and O(log N) ciphertexts.

When the group rekeys (which happens implicitly on every Update, Add, or Remove operation), every member derives the new root secret. The root secret is then run through HKDF-SHA-256 to produce the new epoch's keys: an encryption key, a sender-data key, and an exporter for application use.

The forward secrecy property comes from immediately deleting the old leaf and intermediate keys after the update. The post-compromise property comes from each Update injecting fresh entropy at the leaf, so an attacker who compromised a previous epoch is locked out as soon as the next Update commits.

Epochs, Welcome, and Add

MLS defines a clear notion of "epoch". Every group operation produces a new epoch. Within an epoch, members use a deterministic key schedule to derive per-message keys. Between epochs, the tree changes shape and the root secret changes.

Adding a new member to the group is a special case. Alice generates a Welcome message that contains the new member's path through the tree, encrypted to the new member's public key, plus a snapshot of the current tree shape. The new member processes the Welcome, learns the current epoch's secrets, and immediately becomes a full participant. This is O(log N) work.

Removing a member is symmetric. The group commits an operation that effectively replaces the removed member's leaf with a placeholder, and a path update fresh-keys the path from that leaf to the root. The removed member's old leaf keys are useless because the new path secrets are no longer encrypted to them.

The "everyone updates their key periodically" property is what gives MLS its post-compromise security floor. RFC 9420 recommends every member do an Update every few hours or every certain number of messages, whichever comes first. After enough cycles, the entire tree has fresh keys and any old compromise is forgotten.

The Cipher Suite

RFC 9420 defines several cipher suites for MLS. The default is MLS_128_DHKEMX25519_AES128GCM_SHA256_Ed25519. The components are:

  • HPKE with X25519 as the DH KEM (RFC 9180)
  • AES-128-GCM for application messages
  • HKDF-SHA-256 for key derivation
  • Ed25519 for member signatures

A second suite uses P-256 and AES-128-GCM for FIPS-friendly deployments. A third uses X448 and AES-256-GCM for long-term security headroom.

Notably, RFC 9420 reserved cipher suite numbers for post-quantum variants. Suite 0x0007 in the IANA registry is reserved for ML-KEM (Kyber) variants, and suite 0x0006 is reserved for hybrid X25519+Kyber. The actual definition of those suites is in a separate draft, and as of early 2026 it is still being finalized.

Where the Quantum Risk Is

MLS uses HPKE (RFC 9180), the Hybrid Public Key Encryption standard, as its underlying KEM-and-encrypt primitive. Each Update message encrypts the new path secrets using HPKE to the sibling public keys. HPKE itself supports X25519, P-256, P-384, P-521, X448, and (in extensions) ML-KEM. The MLS cipher suite picks the HPKE KEM mode.

If the cipher suite is X25519-based, an attacker who records the entire group's lifetime can later, with a quantum computer, decrypt every Update path. From every Update, they recover that epoch's root secret. From every root secret, they recover every per-message key. The whole group history becomes readable.

The Ed25519 signatures on group messages are not as critical for confidentiality (they are about authentication, not encryption), but a quantum attacker can forge them retroactively, which is bad for non-repudiation.

The mitigation is the same hybrid pattern: replace the HPKE KEM mode with hybrid X25519+ML-KEM-768, replace Ed25519 with hybrid Ed25519+ML-DSA-65. The IETF draft draft-ietf-mls-pq-postquantum.txt and related work track this.

The cost of hybrid TreeKEM is significant. ML-KEM-768 ciphertexts are 1088 bytes each, and an Update path encrypts O(log N) of them. For a 1000-person group with log2(1000) ~ 10 path nodes, an Update with hybrid is around 10 KB of ciphertext per Update instead of about 320 bytes for X25519 alone. That is a 30x increase in Update size, which adds up over many Updates.

Implementations Available

OpenMLS, an open-source Rust implementation maintained by the MLS WG, is the reference. It supports the standard cipher suites and is the basis for several production deployments.

Cisco Webex was an early adopter and shipped MLS for Webex Meetings in 2023.

Discord uses MLS for end-to-end encrypted voice and video calls in some channels, announced publicly in late 2024.

Signal does not use MLS today; their group encryption is a custom design called Sender Keys plus pairwise envelopes. Whether Signal will adopt MLS is an open question.

Wire announced in 2022 that they were migrating their proprietary group encryption to MLS. The rollout was completed in stages through 2023.

For PQ MLS, OpenMLS has draft branches with hybrid Kyber. Production cipher suites have not yet been finalized at IANA.

How MLS Compares to Megolm and to Sender Keys

The two main alternative designs for group end-to-end encryption are Megolm (used by Matrix) and Sender Keys (used by WhatsApp and Signal groups).

Megolm and Sender Keys both have a single ratcheting key per sender. They scale to large groups by avoiding the per-message asymmetric work that MLS does. They have weaker post-compromise security because key rotations happen on schedule, not on every operation.

MLS pays a higher per-Update cost in exchange for stronger post-compromise security and faster reaction to membership changes. For groups that change often (corporate teams adding and removing people, military units rotating members, hospital shift handoffs), MLS is much better. For groups that are stable and just want efficient bulk messaging, Megolm or Sender Keys are simpler.

Both styles can be made post-quantum. Hybrid MLS replaces the HPKE KEM with hybrid PQ. Hybrid Megolm replaces the Olm transport that delivers the Megolm session key with hybrid PQ Olm. Both inherit the same PQ-AES-256-GCM bulk encryption.

Forward Secrecy in MLS

Forward secrecy in MLS is excellent. Every Update commits new path secrets and immediately deletes the previous ones. Once the deletion happens, no future compromise of the new state reveals the old state.

The "exporter" output of MLS gives applications access to the group's epoch secret indirectly through HKDF, so applications can derive their own purpose-specific keys. RFC 9420 explicitly forbids exposing the raw epoch secret. This is critical for forward secrecy: the application key is derived through HKDF-Expand with a label, so the application can only get a one-way function of the epoch secret, not the secret itself.

Many MLS implementations rotate the epoch on every application message, achieving message-level forward secrecy. This costs more bandwidth (one Commit per message) but makes the post-compromise window as small as one message.

The Federation Question

MLS is fundamentally a centralized protocol. There is a single Delivery Service (DS) that orders all the messages and a single Authentication Service (AS) that verifies identity. The DS does not see plaintext but it sees the message graph and the membership.

Federated MLS is an open research problem. Some proposals have a per-group DS that any participating server can run. Others use distributed agreement to order messages. None has been ratified.

This is why Signal, WhatsApp, iMessage, and Discord can adopt MLS easily — they are centralized — but Matrix and XMPP have not. Federated systems use Megolm or proprietary designs.

QNSQY's Connection to MLS

QNSQY is a post-quantum cryptography tool, not a group messaging protocol, but the design lessons apply directly. QNSQY uses HPKE-style envelopes (KEM + AEAD), the same RFC 9180 HPKE that MLS uses internally. QNSQY's hybrid construction (ML-KEM + X25519, ML-DSA + Ed25519) is essentially the same primitive set MLS uses for its hybrid PQ work. See Hybrid Encryption and ML-KEM Explained.

For organizations that want to layer end-to-end at-rest protection on top of MLS-encrypted messaging, encrypting attachments with QNSQY before sending them through MLS gives a defense-in-depth that is post-quantum at the file level today.

Frequently Asked Questions

Is MLS post-quantum?

Not in its default cipher suite. The PQ cipher suites are reserved at IANA but their definitions are still being finalized in IETF drafts. Expect production hybrid PQ MLS in 2026 to 2027.

How big a group does MLS support?

The protocol is designed for groups up to about 50,000 members. Implementation efficiency varies. Real deployments tested at 10,000-member groups perform well on commodity hardware.

Does MLS need a server?

It needs a Delivery Service to order messages. The DS does not see plaintext but does see metadata: who is in the group, when messages are sent, what size the messages are.

Is MLS used by Signal?

Signal does not use MLS for groups today. Their group design is a custom protocol. They have not announced MLS plans publicly.

What is the difference between MLS and TreeKEM?

TreeKEM is the core data structure. MLS is the full protocol that uses TreeKEM plus HPKE plus a key schedule plus message framing plus signatures. TreeKEM alone is not enough.

Sources

  1. Barnes, R., Beurdouche, B., Robert, R., Millican, J., Omara, E., Cohn-Gordon, K. "The Messaging Layer Security (MLS) Protocol." IETF RFC 9420, July 2023. https://datatracker.ietf.org/doc/html/rfc9420
  2. Barnes, R., Bhargavan, K., Lipp, B., Wood, C. "Hybrid Public Key Encryption." IETF RFC 9180, February 2022. https://datatracker.ietf.org/doc/html/rfc9180
  3. Barnes, R., Beurdouche, B., et al. "The Messaging Layer Security (MLS) Architecture." IETF RFC 9750, January 2024. https://datatracker.ietf.org/doc/html/rfc9750
  4. NIST FIPS 203. "Module-Lattice-Based Key-Encapsulation Mechanism Standard." August 2024. https://csrc.nist.gov/pubs/fips/203/final
  5. NIST FIPS 204. "Module-Lattice-Based Digital Signature Standard." August 2024. https://csrc.nist.gov/pubs/fips/204/final
  6. OpenMLS Project. "OpenMLS Reference Implementation." https://github.com/openmls/openmls
  7. Bhargavan, K., Beurdouche, B., Naldurg, P. "TreeKEM: Asynchronous Decentralized Key Management for Large Dynamic Groups." 2018. https://hal.inria.fr/hal-02425247

Related Articles

Protect Your Data Before Q-Day Arrives

QNSQY's NIST-standardized post-quantum encryption protects files against both current and quantum-era threats.

Try QNSQY