M2 / Mathematical foundations

Hash-based and code-based intuition

FoundationsPractitionerAdvisor

After this lesson you can

  • Explain what 'stateful' means for hash-based signatures (SLH-DSA, LMS, XMSS) and why security collapses completely if key state is reused
  • Explain, in the right frame, why the code-based approach (HQC) offers a different trust model from lattices, and what standardization status it has today

Before thisLattice and LWE intuition

FOUNDATIONS: what are hash-based and code-based signatures and encryption, in one sentence? (read first)

Alongside the lattice-based ML-KEM and ML-DSA you saw in the previous lesson, NIST standardized two completely different mathematical families. Hash-based signatures (SLH-DSA, and LMS/XMSS in the HSM and firmware world) rest their security on a hash function (such as SHA-2 or SHAKE) being hard to invert: not an algebraic problem like lattices, but a much simpler and long-trusted assumption. Code-based encryption (HQC) rests its security on the hardness of “correcting an erroneous message” (error-correcting codes), a family studied since 1978. The hash-based side has a trap: some schemes like LMS/XMSS are stateful, meaning each signing key can be used only once; if the same key is used for two different messages, an attacker can sign any message they like with it. That is why NIST also standardized SLH-DSA, which needs no state. The one sentence to say in a meeting: “NIST did not put all its eggs in one mathematical basket (lattices); the hash-based and code-based families are truly independent backups kept in case something different breaks, but some hash-based variants need careful key-state tracking.”

Mental model

In the previous lesson you saw that the security of ML-KEM and ML-DSA rests on Module-LWE (an algebraic, lattice problem). The two families in this lesson deliberately use a different mathematical foundation: hash-based signatures rest on no number-theoretic or algebraic assumption at all, only on a hash function’s property that “going back from output to input is hard”; code-based encryption rests on the hardness of the decoding problem. This difference matters much more than “which algorithm is faster”: if there is a cryptanalysis breakthrough on the lattice side, the hash-based and code-based families are not affected, because their mathematical foundations are independent.

Hash-based signatures: a simple trust model, but with a trap

Even NIST’s first PQC framing report in 2016 says it plainly: “Hash-based signatures are digital signatures constructed using hash functions. Their security, even against quantum attacks, is well understood.” That is the least contested trust claim among the three PQC families; a hash function’s preimage resistance is a much older and much simpler assumption than the worst-case hardness proofs of lattices.

But the very next part of the same report names the trap: “Many of the more efficient hash-based signature schemes have the drawback that the signer must keep a record of the exact number of previously signed messages, and any error in this record will result in insecurity.” That is what being stateful means: in schemes like LMS and XMSS, the private key is really a large set of “one-time signature” (OTS) keys, and each member of that set can be used only once. In SP 800-208’s own words: “If an attacker were able to obtain digital signatures for two different messages that were created using the same OTS key, then it would become computationally feasible for that attacker to forge signatures on arbitrary messages.” So using an OTS key twice is not a theoretical weakness but a direct path to forgery: from two different messages signed with the same key, an attacker can extract enough to sign any message with that key.

This risk is not abstract: restoring from a backup, cloning a virtual machine, or an unsynchronized multi-signer setup can accidentally use the same OTS key twice. That is why SP 800-208 requires state management to be done in hardware, in a non-exportable way (“This recommendation requires that key and signature generation be performed in hardware cryptographic modules that do not allow secret keying material to be exported, even in encrypted form”). A concrete example: in the NIST-approved XMSS-SHA2_20_256 parameter set the tree height is 20SOURCED, so one key’s total signature budget is 1048576DERIVED; exceeding that number or reusing a leaf directly triggers the forgery risk above.

SLH-DSA: the same family, no state risk

This is exactly why NIST standardized SLH-DSA alongside LMS/XMSS (FIPS 205, FINAL2024-08-13). FIPS 205’s own definition: “This standard specifies the stateless hash-based digital signature algorithm (SLH-DSA)… SLH-DSA is based on SPHINCS+, which was selected for standardization as part of the NIST Post-Quantum Cryptography Standardization process.” SLH-DSA keeps the same hash-based trust model while using a different internal structure (few-time FORS signatures plus a multi-layer hypertree), and removes the need to track state; the price, as you will see in M4, is a larger signature than LMS/XMSS. FIPS 205’s own applicability clause confirms this dual structure: “Either this standard, FIPS 204, FIPS 186-5, or NIST Special Publication 800-208 shall be used”, listing SLH-DSA in the same sentence as SP 800-208 (which defines LMS/XMSS), as alternatives. For a bank’s general-purpose signing infrastructure the right default is SLH-DSA; LMS/XMSS should only be considered for narrow scenarios where the state guarantee can be provided in hardware (such as firmware signing, the use SP 800-208 itself recommends).

Code-based: a third, independent trust model

Beside the lattice and hash-based families, on 11 March 2025 NIST selected HQC (Hamming Quasi-Cyclic) as a backup KEM to ML-KEM. NIST’s stated rationale is clear: “We are announcing the selection of HQC because we want to have a backup standard that is based on a different math approach than ML-KEM.” HQC’s security rests on the hardness of decoding, correcting an erroneous codeword to the right one: a problem family McEliece proposed in 1978, which NIST’s 2016 report describes as having “not been broken since”, and which is completely independent of lattices. The historical drawback of code-based schemes is large key size (NIST’s 2016 report: “most code-based primitives suffer from having very large key sizes”); HQC’s quasi-cyclic structure is an attempt to shrink that size substantially compared with classic McEliece, an engineering compromise similar to how ML-KEM’s module structure shrinks key size (M3).

Status, honestly: HQC is SELECTED2025-03-11 today (September 2026); NIST’s own CSRC project page (as of its 5 August 2026 update) shows that no draft FIPS has been published yet. NIST’s own estimate in the announcement was a draft about a year after selection (around early 2026), with the final standard expected in 2027. That is not the same level of certainty as PKCS#11 v3.2 (a full OASIS Standard) or FIPS 203/204/205 (full FIPS). Telling an architect “HQC is a standard too” carries a different certainty claim from “HQC was selected, there is no draft yet, and tying production decisions to it is premature”.

Three families, three independent trust models: why it matters

What these two lessons (lattice, hash-based and code-based) teach together is not the detail of one algorithm: the four-plus-one algorithms NIST selected (ML-KEM, ML-DSA, FN-DSA, SLH-DSA, and HQC in the future) rest on three independent mathematical assumptions (lattices/Module-LWE, hash function preimage resistance, decoding hardness). The concrete, defensible argument when discussing a bank’s crypto-diversity strategy is: “Our primary choice (ML-KEM/ML-DSA) is lattice-based, but we have real backups resting on independent mathematical foundations (SLH-DSA today, HQC later); if there is an unexpected cryptanalysis breakthrough in lattices, we do not depend on a single mathematical assumption.”

Numbers to know

  • In the NIST-approved XMSS-SHA2_20_256 parameter set the tree height is h=20, so one key's total signature budget is 2^20 = 1,048,576; going past that limit or using the same leaf (OTS key) twice collapses security completely
  • HQC was selected on 11 March 2025 as a backup to ML-KEM; as of September 2026 it is still SELECTED, with no draft FIPS published yet (NIST's own estimate: draft around early 2026, final around 2027)

Lab: Read SP 800-208's own 'state' warning and the XMSS parameter table

[not run] This is a reading and verification exercise, not a runnable command

Requires: internet access. Check your setup

shell
# Open nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-208.pdf and read section 1.2 (p.1)
Recorded output
You will find the sentence "If an attacker were able to obtain digital signatures for two different messages that were created using the same OTS key, then it would become computationally feasible for that attacker to forge signatures on arbitrary messages"
shell
# In the same document find Table 10 (p.16) and look at the h value in the XMSS-SHA2_20_256 row
Recorded output
h=20; that means a budget of 2^20 = 1,048,576 signatures; beyond that you need a different parameter set (e.g. h=60 XMSS^MT) or a completely different key

At the table

How to say this in a bank meeting.

To an executive
Hash-based signatures (LMS/XMSS) make sense only for a specific, narrow use (such as signing long-lived firmware that cannot be updated after production). For the bank's general TLS and PKI infrastructure, SLH-DSA or the lattice-based ML-DSA fit better, because the state management risk is unacceptable in production.
To an architect
If you are considering LMS/XMSS, the signing system must have a state mechanism (inside the HSM, non-exportable) that guarantees it never uses an OTS key twice. Without that guarantee, move to SLH-DSA (stateless, same security family, no state risk).
Objection
“"Hash-based signatures are the oldest, best-understood post-quantum family. Why not just use them and skip lattices?"”
Answer
The security of the hash-based family really is well understood (NIST has said so since 2016), but LMS/XMSS themselves require state, and a state management error (restoring from backup, cloning a VM, bad synchronization) breaks the key completely. That is why NIST recommended them for narrow, controlled scenarios (firmware signing) rather than general use, and standardized the stateless SLH-DSA for general purposes. Being 'the oldest' does not mean 'use it everywhere'.

Sources

Checkpoint

Answer first, then compare with the model answer and score yourself against the rubric. Saved in this browser only.

  1. 01Recall

    What does 'stateful' concretely mean in a hash-based signature scheme (LMS/XMSS), and why is losing or repeating state a disaster?

  2. 02Recall

    What standardization status does HQC have today (September 2026), and what does that status mean?

  3. 03Scenario

    A team plans to use XMSS-SHA2_20_256 to sign the firmware of an IoT device and expects 2 million updates over the device's lifetime. What do you tell them?

  4. 04Hostile

    An architect says 'Code-based cryptography is a math problem just like lattices; the difference is unimportant.' How do you answer by distinguishing the trust models of the three families (lattice, hash-based, code-based)?

Project linkContributes to the crypto-diversity and algorithm selection section of Project 4 (capstone), as the rationale for when to prefer each of the three trust models (lattice, hash-based, code-based).