- The thesis
- Position in the tower
- Competitive landscape
- Spreadsheets are a source, not a rival
- What a typed model adds
- Applications
- Model overview
- Relation to other Nasdanika work
- Resources
An Ecore micro-model of the physical data estate: systems of record, datasets, record layouts, fields - and the mapping from a business vocabulary term to the fields that actually hold it.
The name is deliberate. System of Record names the role a system plays - the authoritative holder of a class of records - not a technology. Copybooks are COBOL; the same mainframe also carries PL/I includes, Assembler DSECTs, IMS segments, VSAM layouts and DB2 tables, and the same role is played off-mainframe by Cassandra column families and JSON documents. RecordLayout and Field are therefore abstract; Copybook/DataItem and Table/Column are the first two specializations, and others join as needed without disturbing anything above them.
The thesis
The interesting question about a bank is not what its data model is. It is where the same concept lives, in how many shapes, and who can tell you.
Take one term from the canonical bank model: Transaction.amount, an EBigDecimal. In the estate it is not one thing:
- in the deposits system, a COBOL data item -
05 TRAN-AMT PIC S9(9)V99 COMP-3- packed decimal, implied scale 2, sign in the low-order nibble; - in the credit card system, a
DECIMAL(11,2)column with a separate sign indicator and credits stored positive; - in the auto loan system acquired three years ago, an integer of minor units with the scale carried in a companion currency field.
None of those is wrong. All three are Transaction.amount. VocabularyMapping exists for exactly this shape: many-to-many and context-qualified - a term, a set of fields, a context (product line, line of business, era of the estate), and transformation notes covering scaling, sign handling, code page and enumeration recoding.
That is the whole wedge. Everything else in the model is the structure needed to make that statement precisely, review it, own it, and generate from it.
Two things this model deliberately is not. It is not a transformation engine - transformation is Markdown for humans to review, and its executable counterpart is NSML, not a hidden expression language here. And it is not a replacement for a source metamodel: richer SQL structure belongs to the SQL model, richer COBOL structure to a COBOL parser. This is the lineage-and-meaning level view, at the resolution at which a human can actually assert something and be held to it.
Position in the tower
The aspect spine runs nxcore < role < iam < seal < lifecycle < decision analysis < decision binding < governance < work < architecture. This model branches off architecture: it is a structural estate model, so it sits at the lowest floor consistent with its own dependencies, one step above the elements it describes.
Three consequences, and they are the point of the arrangement.
A system of record is an architecture element. SystemOfRecord extends architecture’s Element rather than pointing at one by URI. The deposits system is not described twice - once in the architecture model as a software system and again here as a data holder. It is the same object seen through a different lens, exactly as the threat model makes an Asset an Element instead of maintaining a parallel YAML description of the same estate.
Every element down to the field inherits the whole spine. Because architecture’s AbstractElement is Workable, and Workable is Governed, and governance descends through decision binding, decision analysis, lifecycle, seal, IAM and role, a single Column arrives already carrying: an owner and a steward (role), access control (IAM), attestation signatures (seal), a lifecycle stage (lifecycle), the decision and rationale that put it there (decision analysis and binding), applied controls and evidence (governance), and open work (work). None of that is declared in this model. Declaring it here would have been the mistake.
Relationships between estate elements are reified and can be tunnelled. A field-to-field feed is an architecture Relationship, and carries gives the layered view for free: the logical statement “cards TRAN-AMT feeds the risk mart TXN_AMT” rides on the physical statement “nightly extract to landing zone”, and either layer can be queried without inventing a second relationship model.
The placement rule is worth naming because it is tempting to get wrong. This model does not branch off threat, even though data protection is obviously in scope, because it has no typed need for weakness containment. Personal-data classification arrives as a governance Control applied to a Field - governance is below architecture, so the capability is inherited rather than added. If a typed need appears later (field-level weaknesses, say), the model moves up a floor; until then it stays where its dependencies put it.
Competitive landscape
The actual incumbent: Word, Excel, Confluence and SharePoint. A data dictionary tab per system, a Confluence page per interface, a mapping spreadsheet per project, and a Visio diagram nobody has opened since the author left. This is not a straw man - it is what the practitioner literature still recommends. Guidance on GDPR Article 30 records openly advises that “a Google Sheet or Excel file backed by the ICO template is perfectly adequate”, and prescribes building the map by interviewing process owners across engineering, HR, marketing and support, because “each know things no single document does”. That is an accurate description of how enterprise field mapping is done today. It is also a description of a method whose output is stale on the day it is filed, cannot be diffed, and cannot be queried.
The response, though, is not to ask anyone to stop using Excel. It is to stop treating the workbook as the authority - because the workbook loads. See Spreadsheets are a source, not a rival below.
Data catalogs and governance platforms. Collibra, Alation, Informatica, Atlan, Ataccama, data.world, and the open-source line - OpenMetadata, DataHub, Amundsen, Egeria. Genuinely good at the things this model does not attempt: automated harvesting, stewardship workflow, search, access requests, popularity signals. Two gaps. First, the estate lives in the platform, and the crosswalk to architecture, risk register, control catalog and backlog is a manual export or absent - the same critique the tower makes of EA and GRC platforms. Second, and more fundamental, their “business glossary” is a set of prose terms in the same tool, not a typed model you can generate code from. The industry’s own retrospective on why glossaries become shelfware lands precisely here: adoption drops when glossary terms are disconnected from actual datasets, and analysts do not consult a reference that does not connect to where they work. This model’s term points at a real EClassifier in a real model - the bank model - and the mapping to a real Field is a typed reference, not a hyperlink in a wiki.
Lineage specialists. Solidatus, IBM MANTA, Octopai, OpenLineage/Marquez, dbt docs, Precisely, Zengines. The strongest camp technically, and the closest on the mainframe: Solidatus advertises field-level lineage extracted from COBOL, JCL and 50+ mainframe languages, MANTA traces issues through COBOL applications into downstream analytics, Precisely Connect moves copybook-defined data with EBCDIC conversion, and open-source Cobrix parses REDEFINES, OCCURS DEPENDING ON and variable-length structures faithfully. They are sources, not competitors: they answer where did this column come from, mechanically, from code. They do not answer which business concept is this, under which product line, and who is accountable for the answer - that half is asserted by a human, must be reviewed like code, and is what this model is for. Their output is also a graph inside a vendor UI; this one is a file in Git.
Data modeling tools. erwin Data Modeler, ER/Studio, SAP PowerDesigner, SqlDBM, Hackolade. The nearest prior art in intent - logical and physical models with forward and reverse engineering, which is canonical modeling as actually practised. Three costs: a proprietary repository, a diagram-centric authoring model, and the built-in assumption that one logical model maps to one physical model. The moment Transaction.amount maps to three differently scaled fields under three product lines, the tool’s mapping facility is the wrong shape and the work moves to a spreadsheet.
Mainframe discovery and modernization. IBM ADDI and watsonx Code Assistant for Z, OpenText (formerly Micro Focus) Enterprise Analyzer, CAST Imaging, Hypercubic. They parse copybooks, JCL and DB2 DDL and build call graphs and dependency trees - again, sources. They produce the layouts this model imports; they stop short of the business vocabulary and have nothing to say about ownership, controls or work.
Semantic and metrics layers. dbt Semantic Layer, Cube, AtScale, LookML, Malloy. The modern rediscovery of the canonical model, and a good one - but scoped to the analytical estate and to SQL by construction. The operational systems of record where the record actually lives (copybooks, IMS segments, VSAM) are out of scope on day one, which is where the hard mappings are.
Industry canonical models and metadata standards. BIAN, FIBO and ISO 20022 in banking, ACORD in insurance, and the older lineage of OMG’s CWM and ISO/IEC 11179 metadata registries. These are the canonical vocabularies, and they are sold as documents and spreadsheets. They are export targets and inputs: a published industry model loads as the canonical model that term points at, and adopting one becomes authoring rather than a metamodel change - the same posture the tower takes toward C4, ArchiMate and control frameworks.
And the honest observation underneath all of it. In most estates there is no formal canonical data model at all. The closest thing is whatever the newest system’s API happens to return, and the mapping between a screen and the systems behind it exists only as tribal knowledge: a UX designer draws fields in Figma, and someone works out where each value comes from by asking two or three teams. The handoff literature confirms what goes missing at that boundary - data requirements, “what API endpoints, what data structure”, are named among the elements typically absent from a design handoff, alongside edge cases, empty states and error states. The competitor is not a better tool. It is a conversation that has to be repeated every time somebody new asks.
Spreadsheets are a source, not a rival
The critique above is of the spreadsheet as the system of record for the system of record - the place the authoritative answer lives, unversioned and unqueryable. It is not a critique of spreadsheets as an editing surface, and the tower’s standing posture applies here as everywhere: loaders treat existing artifacts as source rather than export.
A workbook named by convention - deposits.sor.xlsx - loads as SOR model contents. Nasdanika’s resource contents filters read filename qualifiers right to left: the rightmost qualifier is the source format, and each qualifier to its left is a transformation stage over the output of the previous one. So .xlsx loads the workbook and .sor transforms workbook contents into this model. ResourceSet.getResource(uri) hands back typed SystemOfRecords, Tables, Columns and VocabularyMappings, and nothing downstream knows or cares that the source was a spreadsheet. Chains compose the way the rest of the tower composes - estate.html.sor.xlsx runs workbook → model → site in one resolution.
The transformation itself is either an NSML rule set - declarative, diagrammable, and reviewable as a model in its own right - or a Java ResourceContentsFilter. Filters implement save() as well as load(), so a round trip is available where it earns its keep: the steward keeps maintaining the workbook they already maintain, and the model is regenerated from it on every build.
That is the move the incumbent stack cannot make at any price. The sheet stops being the record and becomes an input to the record. A column list on a tab is a perfectly good way to enter two hundred Columns quickly; a mapping tab with term / system / field / context / transformation columns is a perfectly good way for an analyst who will never open Xcore to state a VocabularyMapping. Loaded elements carry markers back to workbook, sheet and row, so provenance survives the lift and the next refresh is a model diff rather than a fresh spreadsheet. The same applies to the other incumbents: Confluence pages, Word documents and Visio diagrams are loadable sources under the same mechanism, and the Excel and Drawio models are where those formats already live in the tower.
One boundary stays exactly where it was. A flat field-per-row sheet still cannot express REDEFINES or OCCURS DEPENDING ON - that is a property of the sheet, not of the loader, and no naming convention repairs it. Copybook layouts therefore come from copybook members and from the extraction tools that already parse them; spreadsheets are the right source for flat column lists, ownership columns and mapping tabs, which is most of what is in them anyway. Excel for the entry, the model for the record.
What a typed model adds
One term, many fields, qualified by context. The relationship is many-to-many by construction, not by convention. “Which physical fields realize Transaction.amount, and under which product line” is a query; on the incumbent stack it is three meetings.
Copybooks survive the round trip. DataItem carries level, picture, usage, occurs, dependingOn and a redefines reference to a sibling item. That last one is the concrete answer to “why not a spreadsheet”: a REDEFINES means two names occupy the same storage, and a field-per-row table cannot represent it without lying. OCCURS DEPENDING ON has the same property. A model that quietly flattens them produces a document that is confidently wrong about the record you are trying to migrate.
Ownership, controls and work land on the field itself. Because of the tower, “who owns TRAN-AMT” is a role engagement, “is this field personal data” is a governance ControlApplication with evidence, “when is this layout retired” is a lifecycle stage, and “what is open against it” is contained work. On the incumbent stack these are four artifacts in four tools, each with its own copy of the field list, already inconsistent.
Traceability runs in both directions from the canonical model. Downward, VocabularyMapping reaches the physical field. Upward, the UI model’s ValueBinding reaches the screen element bound to the same term. Which makes “if the cards system changes the scale of TRAN-AMT, which screens, reports and interfaces are affected” a traversal rather than an archaeological dig - and makes the Figma-to-backend conversation a generated document instead of a recurring meeting.
Regulatory lineage falls out as a view. The ECB’s May 2024 Guide on effective risk data aggregation and risk reporting expects complete and up-to-date data lineage at the data attribute level, from data capture through to final reporting, and RDARR deficiencies sit at the top of the ECB’s supervisory priorities for 2025–2027 - eleven years after BCBS 239 was published in 2013 with a 2016 target, with implementation still assessed as unsatisfactory. Extraction tools produce the mechanical half of that record. The semantic half - this attribute is that risk metric, in this context, with this transformation - is asserted, and this is where the assertion lives, reviewable in a pull request.
The model can say “we do not know”. A term with no mapping is a visible hole; a mapping with an empty transformation is a visible question. A spreadsheet with a blank cell is indistinguishable from a spreadsheet nobody finished.
Applications
Educational - understanding backend systems. The most under-served audience is the one that most needs this: analysts, designers, product managers and new engineers who have to reason about systems they will never log into. The domain is normally taught as folklore. A metamodel is the alternative: an Estate contains SystemOfRecords, a system contains Datasets, a dataset contains RecordLayouts, a layout contains Fields, and a VocabularyMapping ties a business term to the fields that hold it. The COBOL specialization doubles as a syllabus - level numbers, PIC clauses, USAGE, COMP-3, OCCURS, REDEFINES - with the semantics drawn from public IBM Enterprise COBOL documentation. The instructive exercise is instantiation: take one term, map it across three systems, and write the transformation notes. What you cannot fill in honestly is the lesson.
Documentation generation. A browsable site per estate: a page per system, per layout, per field, per canonical term, with the mapping rendered from both ends. Generated by the same stack that built this page, from a source that is diffable and reviewable.
Association of ownership. Data stewardship stops being a spreadsheet of names. Owners are role engagements on elements that also carry stage, controls and work, so “unowned fields in scope for risk reporting” is a query and a stewardship report is a view.
Controls, privacy and audit. Personal-data classification, retention, encryption at rest and masking are governance controls applied to fields, with evidence attached. A GDPR Article 30 record, a data protection impact assessment’s data inventory, and a BCBS 239 attribute lineage pack become three views of one model rather than three documents that disagree.
Work: migrations, decommissioning and modernization. “What maps to the system we are retiring, and what breaks when it goes” is a traversal. Target-state mappings sit next to current-state mappings under a different context, so a dual-run reconciliation plan is generated rather than assembled. The decisions taken along the way - which system becomes the record for which term - are recorded as bound variation points, so the rationale outlives the programme.
Impact analysis and change management. Field-level change requests answered against the model instead of against memory; the blast radius of a scale change, a code-page change or an enumeration recoding is enumerable.
Traceability to UI and other artifacts. Screens via the UI model, reports, interfaces and APIs, test data sets, and the requirements that constrain them - all anchored on the same canonical term, so the trace from a field on a screen to the packed-decimal item behind it is a path rather than a project.
Grounding for agents and assistants. “Where does the customer’s available balance come from, and what does it mean in the cards context” is the question every internal assistant gets asked and answers badly. A typed estate model is the retrieval target that makes the answer checkable - and an agent that proposes a mapping produces a reviewable diff rather than a confident paragraph.
Interoperability and adoption without a migration. Loaders from copybook members, DDL, extraction tools’ exports and - the one that decides whether anyone adopts this - the workbooks the organization already keeps, via the *.sor.xlsx convention above. Emitters toward catalog platforms and lineage formats. Day one does not require anybody to stop using their spreadsheet; it requires the spreadsheet to be named and shaped so a build can read it.
Model overview
| Area | Types |
|---|---|
| Base | SorElement (architecture Element, hence Workable, Governed, staged, access-controlled, engageable) |
| Estate | Estate, SystemOfRecord (platform), Dataset |
| Layouts and fields | RecordLayout (source, abstract), Field (abstract) |
| COBOL | Copybook, DataItem (level, picture, usage, occurs, dependingOn, redefines), Usage |
| Relational / columnar | Table, Column (dataType, nullable) |
| Mapping | VocabularyMapping (term, fields, context, transformation) |
| Inherited | nxcore documentation and markers, role engagements, IAM subjects, seal signatures, lifecycle stages, decision variation points and bindings, governance controls and evidence, work items, architecture elements and reified relationships |
Next specializations, in the order the estate usually demands them: PL/I includes, IMS segments, VSAM layouts, Assembler DSECTs, JSON and Avro schemas, Cassandra column families. Each is a RecordLayout/Field pair; nothing above the abstraction changes.
Relation to other Nasdanika work
The canonical model. Bank is the worked example - Customer, Account, Statement, Transaction - and any Ecore model can play the role. The point of using a model rather than a glossary is that the canonical side is itself generatable, diffable and typed.
Transformations. NSML is the executable counterpart of transformation: a declarative match-and-transform language over Ecore models, diagrammable and reviewable, so a mapping can graduate from prose to something that runs.
Upward traceability. The UI model binds screen elements to data; the requirements model states what must be true of them; the threat model treats data stores as assets and flows as crossings of trust boundaries.
Tooling. Models are authored in Groovy DSL or as XMI/YAML/JSON, entered in Excel through the Excel model and resource contents filters, diagrammed with the Drawio model, and documented with the generation stack that built this site.
Nasdanika Models