description: "Eight unresolved challenges (from measurement validation to the standardization of vendor-neutral assurance formats) define the operational workload …"
Chapter 18 · The Research Agenda
What we discuss in this chapter: Eight unresolved challenges (from measurement validation to the standardization of vendor-neutral assurance formats) define the operational workload of this category. Divided into current status, completion criterion, and collaboration format, they form an open invitation to shape the future.
Your leverage as a decision-maker: A discipline that documents its open research questions establishes a serious research program. Pilot partnerships ensure you have a direct influence on future standards, rather than merely purchasing off-the-shelf software. The eight challenges also serve as an in-depth due diligence checklist.
18.1 Agenda Instead of Roadmap
A classic software roadmap promises ambitious features on fixed dates; a scientific agenda names uncomfortable questions to which no one currently has a definitive answer.
This chapter deliberately chooses the second path. Our methodical role model stems once again from process science: The Process Mining Manifesto from 2011 identified eleven unresolved challenges, giving the then-young research field its outstanding credibility and momentum precisely through this transparency. The eight points in this chapter mirror the eight principles of our own manifesto: every strategic commitment corresponds to a concrete operational work site. If you as a decision-maker wish to rigorously evaluate whether this new category can deliver on its promises in your organization, you should measure progress against this exact list.
18.2 The Eight Challenges
H1 · Measurement Validation: While the comprehensive metrics catalog from Chapter 10 is theoretically defined with precision, it has so far been empirically tested on too few real organizations. Although systematic surveys in practical cases A through D are firmly planned, all industry-specific baseline values currently remain open. The current status is limited to the methodical formulation of indicators. The completion criterion is fulfilled only once reliable baseline value ranges from at least ten different organizations are published. Structured baseline surveys with fully anonymized evaluations serve as an effective collaboration format.
H2 · Causality of the Core Hypothesis: The assumption that relational decision quality under turbulent change shapes long-term organizational performance more strongly than pure operational process efficiency is highly plausible, but scientifically unproven. Currently, this core thesis exists merely as a well-founded hypothesis model with derived indicators. As a completion criterion, we require at least one empirical longitudinal or comparative study with proper control of potential confounding variables. The ideal collaboration format here is a joint study design with a renowned university chair.
H3 · Extraction Quality and Human-in-the-Loop: To date, there is no systematic, scientifically published measurement assessing how precisely formal Claims are automatically extracted from unstructured texts and at which interfaces human validation remains indispensable. The current status includes valuable project-related empirical values from pilot applications, but no valid evaluation yet. The completion criterion requires the publication of quantitative precision and recall metrics per source type (e.g., PDF manuals, wiki pages, ERP master data). Manually annotated benchmark corpora from real, anonymized document sets serve as the collaboration format.
H4 · Scaling of Consolidation: It remains to be empirically demonstrated whether the developed conflict semantics functions stably even when Claim bases expand to six figures across dozens of heterogeneous source systems. The current status is limited to the medium project sizes of practical cases A through D. The completion criterion demands proven, long-term production operation beyond this scale without the risk of an organizational review collapse or uncontrolled cascade conflicts. The collaboration format consists of scaling partnerships in complex enterprise environments.
H5 · Governance Patterns for Co-Determination and Data Protection: Although the theoretical protection principles are clearly articulated in Chapters 14 and 15, universally deployable template works agreements for this novel system class are still lacking in practice. The status is limited to project-specific agreements. The completion criterion is a publishable template agreement battle-tested in daily operations, co-developed with works councils, data protection officers, and supervisory authorities. The collaboration format is joint legal and organizational template development with pioneer companies.
H6 · Standardization of Assurance Formats: Assurance evidence from the knowledge base must be vendor-neutrally exchangeable between different software implementations; this is the direct consequence of our manifesto principle and machine-readable assurance standards like OSCAL (Chapter 14). The status shows initial proposals submitted to relevant industry and standardization committees. The completion criterion is proof that at least two independent software implementations fully support the same machine-readable exchange format. The collaboration format is active participation in cross-vendor standardization committees.
H7 · Sovereignty Classification in Practice: Whether and in what form the developed CADA classification finds its way into procurement reality remains an open practical question, as corresponding regulatory proposals and rules are still in political proceedings. The status is pure observation and expert contribution. The completion criterion is the appearance of CADA tier references in official public or regulated tenders. The appropriate collaboration format encompasses continuous monitoring and commentary work in leading industry associations.
H8 · Validation of the Maturity Model: The four developmental stages from Chapter 12 are pragmatically constructed, but not yet scientifically validated. While the digital self-assessment is live, the underlying database is still slim. The status is a prototype. As a completion criterion, we require valid distribution and trajectory data from at least ten organizations, including an empirical inter-rater evaluation of the diagnostic questions. The collaboration format consists of the anonymized evaluation of self-assessment data.
18.3 What We Contribute – and What We Hope For
The contribution of the authors and aiio GmbH can be precisely quantified: we provide the reference architecture developed in this book as an open foundation for discussion, release the developed metrics profiles for free use, actively drive standardization work in committees, and stand ready to supply anonymized project data for challenges H1, H3, and H8. Our request to you and the broader community is articulated just as concretely. We seek ambitious organizations that establish reliable baselines before implementing software; academic chairs that make Hypothesis H2 a research priority; and forward-thinking competitors who take Challenge H6 seriously and collaborate on a neutral exchange format. A sustainable category is never born from a single book. It emerges only when enough actors systematically work through the research agenda of this chapter – very gladly with the objective of empirically refuting individual theses of this book.
💡 What We Discussed
This chapter presented a scientific research agenda that precisely identifies eight central challenges of the new category.
From empirical measurement validation and automated extraction quality to neutral assurance standards, these work sites demand close collaboration between practice and research.
By sharing our own data and architecture models, we lay the foundation for a reliable ecosystem in which you can actively participate.
How this collective development work could unfold over the coming years is illustrated by the methodical analysis of future scenarios in the concluding chapter.
