What the problem looks like
You’ve been on the platform for a few years. TEST and PROD were once mirrors of each other — built from the same templates, promoted through the same release pattern. Time passes. Someone fixes a calculation in PROD because the close was running long and didn’t have time to round-trip through TEST. Someone adds a member to a dimension in TEST to support a model that never ships to PROD. Someone refreshes the security model in one environment without the other. A vendor changes a default.
None of these events feel significant in isolation. None of them would fail an audit on their own. Together, over a few years of normal operation, they produce two environments that are no longer comparable. The schema of the data — what shapes the platform expects, what hierarchies roll up where, what calculations fire on which scenarios — has drifted. We call this a schema-version mismatch.
Schema-version mismatch is invisible until something forces a direct comparison. Then it’s suddenly the only thing on the page.
When it bites
Three events surface it:
- A migration. A new module gets added, or the platform is being moved to a different pod, or a major version upgrade is rolling through. The migration tooling assumes TEST and PROD are comparable. When they aren’t, the migration fails in opaque ways — an LCM (Lifecycle Management) export from one environment cannot be imported cleanly into the other. The error messages point at things that look like permissions or connectivity issues. They are not.
- A close-cycle audit. Auditors ask you to reproduce a specific number from the close. You walk them through PROD. They ask you to confirm the same calculation runs the same way in TEST. It doesn’t. Now you’re explaining why your two environments are different, and the audit shifts from sampling to investigation.
- A rescue. The previous implementer left and you’re bringing in a new team. The new team needs to understand the platform before they can change anything. The handoff documentation is for TEST. The platform that actually runs the business is PROD. Every assumption in the documentation has to be re-verified against reality, and reality has drifted.
In the engagement this note is drawn from, the trigger was the first one: a Crown energy producer was moving toward an upgrade and discovered that the TEST environment they’d been using to validate the upgrade was no longer a faithful representation of PROD. The upgrade plan that had been signed off was structurally sound. The environment underneath it wasn’t.
The diagnostic
EPM Cloud doesn’t ship a diff tool. There is no one-button comparison that tells you where TEST and PROD have diverged. You have to build the comparison yourself, and you have to build it carefully, because the dimensions you’re comparing are not flat.
The pattern we use is layered. Take it in this order:
- Export both environments via LCM. LCM snapshots are XML. They are not optimized for diffing, but they are deterministic enough that a structural comparison surfaces the obvious drift first.
- Diff the metadata first, not the calculations. Dimension members, alias tables, attribute associations. Calculations depend on metadata; if the metadata is different, you cannot meaningfully compare what the calculations do. Settle the metadata before you look at code.
- Diff the security model second. Member access, application access, artifact access. A permission that was granted in PROD as a one-time fix and never replicated to TEST is the second-most-common cause of divergence we see.
- Then diff the business rules. Calculation scripts, rulesets, runtime prompts. By this point you’ve already eliminated the differences thatcauserule differences; what’s left is real rule drift.
- Finally, diff the data. Specific intersections, run the same calculations, see what produces different numbers. By this point the “why” of any difference is traceable.
Each step has tooling around it. EPM Automate exports the snapshots; Python (or any structured-diff tool) compares them; human judgment decides which differences are intentional and which are drift. The discipline is doing the layers in order, not skipping ahead to the data and trying to back-explain.
The recovery
Once you know where the divergence is, you have three options for each one: bring TEST forward to match PROD, bring PROD back to match TEST, or accept the divergence and document it as intentional. The first option is the most common. The second is almost never right, because PROD is the system of record. The third is acceptable only when the divergence is load-bearing — a feature in TEST that was never shipped to PROD on purpose, for example.
In the Crown energy engagement, the recovery was a focused piece of work across the diagnostic and remediation steps. The upgrade went ahead on a revised schedule, against an environment pair that was actually comparable, and the close cycle that followed ran clean.
How to prevent it
Schema-version mismatch is a discipline problem before it’s a tooling problem. The two practices that catch it early:
- Every change goes through TEST first, every time, no exceptions for emergencies. The temptation to fix PROD in a hurry is the single biggest source of divergence. Manage the cost of doing TEST-first by making TEST-first cheap, not by skipping it.
- Snapshot-based regression check on a cadence. Once a month, take LCM snapshots of both environments, run the layered diff, surface anything new. This is the same tooling you’d use in a rescue — running it preemptively catches the drift before it accumulates.
Most platforms with a multi-year history have some amount of schema drift in them. The question is whether you find out about it on your terms or on an auditor’s.