M13 / Migration program

The cost model: methodology and the corpus's own provenance lesson

Advisor

After this lesson you can

  • Explain the methodology behind the cost model tool (bank parameters in, a range out, every assumption listed and editable, tagged ESTIMATED)
  • Examine the corpus's own CAPEX contradiction (four different numbers and no single 'true' total) as a provenance lesson

Before thisPrioritization and sequencing: three horizons

Mental model

In the previous lesson you built a three-horizon roadmap. Each horizon has a cost, and that is exactly this lesson’s subject: how to estimate the cost of migration, and seeing how this course’s own source material answered that question wrongly.

The corpus’s own contradiction: four numbers, zero agreement

This course’s raw source material (the Day 14/15 documents) has four different figures for the 3-year migration CAPEX, and none agrees with any other:

  1. Day 14, a partial-scope calculation: EUR 1.2M - 2.5M
  2. Day 14, in the same document, a full itemized breakdown: EUR 2.05M - 3.9M (all components are unsourced round numbers, and the figure is wrongly described as “conservative”, while the DigiCert survey it is compared with measures only a partly overlapping scope)
  3. Day 14, an executive summary rounding: EUR 2-4M
  4. In a YAML export, with false precision: capex_eur_low_m: 2.05, capex_eur_high_m: 3.9

On top of that, Day 15 derives the cost of the discovery phase as a percentage of Day 14’s, but along two different, contradictory paths: in one place it says “~25-30%” to get EUR 650K-1.35M, in another “40-55%” to get EUR 1.4M-2.2M. The second percentage does not even match its own stated base (40-55% of EUR 2.05-3.9M): the correct arithmetic gives ~EUR 0.82-2.15M, not 1.4-2.2M.

This is a provenance lesson: why we use none of them

None of these four figures entered this course’s numbers.json as SOURCED or MEASURED, because none rests on a real source (an audit report, a real bank project, a published methodology); all the itemized components are unsourced, and the two derivation methods contradict each other. The right way to fix this is not to pick “whichever looks more right” (recall that in M6 a BSI date appeared taken from the wrong document and matched to the wrong year, and that the corrected version is documented in a note on 2031SOURCED; picking a figure that “looks more reasonable” is also a kind of fabrication), but to accept none of them as true.

Instead, the cost model tool works on this principle: bank parameters (such as number of systems, certificate inventory, staff capacity) are taken as input, each assumption (for example “X person-hours per certificate”) is listed explicitly and is editable, and the output is a range always presented with an ESTIMATED tag, never as a “true” total. DigiCert’s own survey data is third-party data this course could not verify independently (so it cannot enter numbers.json as SOURCED yet); if it can be verified, it may be one of several data points informing the default values, but on its own it is not a reference that guarantees the “true total”.

The scope here: methodology, not the tool

This lesson teaches not the tool itself but the methodology behind it: (1) break down the cost components (inventory and discovery, certificate renewal, HSM and infrastructure upgrades, staff training, testing and validation), (2) write an assumption for each component and state its source (a SOURCED data point, or an ESTIMATED guess), (3) produce a range, not a single point estimate, (4) keep every assumption transparent, so that it can change when the bank enters its own parameters. The tool makes this methodology interactive; this lesson teaches how to set those assumptions defensibly for an atypical bank before you use it, because running the tool is not enough: you need to be able to defend the range it produces against your own bank’s reality.

Numbers to know

  • The corpus's (the course's raw source material) own CAPEX claims CONTRADICT each other, and none could enter numbers.json as SOURCED: on Day 14 there are three different '3-year totals' (partial scope EUR 1.2-2.5M; full breakdown EUR 2.05-3.9M; executive summary rounding EUR 2-4M) plus figures of 2.05/3.9 with false precision in a YAML export; on top of that, Day 15 produces two different discovery cost percentages that do not agree either (~25-30% → EUR 650K-1.35M; 40-55% → EUR 1.4-2.2M, and the 40-55% figure does not even match the arithmetic on its own Day 14 base)
  • No CAPEX figure in this lesson is presented as SOURCED/MEASURED with a ProvenanceBadge; all are told in plain text as EVIDENCE of the corpus's own internal inconsistency, clearly marked as 'data this course rejects'

Lab: Build an assumption table for your own bank (by hand, without the tool)

[not run] This lesson applies the methodology by hand; the cost model tool makes it interactive

Requires: . Check your setup

shell
# Build an assumption table from scratch: list each input, such as number of systems, number of certificates, number of HSMs, and estimated person-hours per system, on its own row, and write down its source (this course's DigiCert data, or your own estimate)
Recorded output
A table: each CAPEX component, its default value, its source (SOURCED/ESTIMATED), and why that value was chosen

At the table

How to say this in a bank meeting.

To an executive
We do not present migration cost as a single number, because there is no single correct number: the cost varies with your bank's number of systems, certificate inventory and current crypto-agility. Instead we offer a range tool where you can see and change every assumption.
To an architect
This course's own raw source material had four different CAPEX figures that did not agree with each other; the way to fix that was not to pick 'which one is right' but to accept none of them as true and move to a transparent, assumption-based model.
Objection
“"I just want one number. This range and assumptions business is too complicated."”
Answer
Giving one number is exactly the mistake this course's source material made: it produced four different 'single numbers', none agreeing with any other, and none showing which assumptions it rested on. A range plus visible assumptions is more honest and more useful than false precision, because you can adjust it to your own bank's real parameters.

Sources

Checkpoint

Answer first, then compare with the model answer and score yourself against the rubric. Saved in this browser only.

  1. 01Recall

    How many different '3-year CAPEX totals' are there in the corpus's Day 14, and do they agree with each other?

  2. 02Recall

    Why are Day 15's two different discovery cost percentages also mathematically inconsistent?

  3. 03Scenario

    A board member reads the 'EUR 2-4M' figure in a presentation as a 'firm budget'. How do you correct this misunderstanding?

  4. 04Hostile

    An auditor asks 'Does this course have its own CAPEX numbers, or is it all estimates?' Explain why you did not use the corpus's four contradictory figures and what you propose instead.

Project linkContributes to the cost section of Project 4's (capstone) migration roadmap, both as a methodology and as the discipline of 'a transparent range instead of one number'.