Bill Lynch, our co-founder and CTO, has been named to the CDISC USDM Governance Group (UGG), the group responsible for guiding the evolution of the Unified Study Definition Model. CDISC announced the new cohort, and we are glad to be part of it.
We build on USDM every day, which gives us a direct stake in keeping it usable, well-reasoned, and well-governed. That is why this group matters to us, and why we put our name forward. We also think we have a bit of unique take on how the standard is used, which we'll explain below. But first, a bit of background on what the USDM is and what the UGG does.
A quick USDM refresher
The Unified Study Definition Model is one of the CDISC standards. In plain terms, it is a single, machine-readable way to represent a clinical trial definition, the protocol and its structure, as data rather than prose. For decades (and still today), the protocol has been shared as an unstructured document, usually as a PDF. This causes predictable downstream problems like misinterpretation, copy-and-paste errors, and a lack of programmatic interoperability.
That shift from document to data is the foundation of the industry's Digital Data Flow (DDF) vision. If a protocol can be expressed as structured, conformant data, then the systems downstream of it (randomization, EDC, consent, and more) can consume that definition directly instead of re-keying it from a PDF. The USDM is how a protocol stops being a document and starts being data.
What the USDM Governance Group does
The UGG provides oversight, decision-making, and strategic direction for the USDM. Its job is to make sure changes to the standard are well-reasoned and technically sound, aligned with the needs of sponsors, vendors, and technology providers, and supportive of the broader Digital Data Flow vision and CDISC's commitment to interoperability.
It is a small, community-driven group: a handful of volunteer representatives working alongside the CDISC leads, each serving a one-year term. For a standard still maturing in production, that kind of community governance matters. The decisions made now about what the model should represent, and how, will shape what the entire ecosystem can build on top of it.
Why HumanTrue is involved
We think we bring a useful and slightly unusual vantage point to that work. We use the spec itself to inform our use of Large Language Models to reconstruct structured study definitions from unstructured protocol documents. That means we have had to understand the spec not just as a schema to populate, but as a design document that informs how we build our AI pipeline and how we interpret the content of a protocol. We also use the spec to inform user interface decisions about how to present the reconstructed data back to users, and how to make it actionable for downstream systems.
USDM was largely designed for authoring protocols from scratch. It assumes a born-digital workflow, where the model is the source of truth and the structured definition comes first. We come at it from the opposite direction. We treat USDM as the engineering target that informs an AI pipeline, reconstructing structured study definitions from unstructured protocol documents that were never written with a data model in mind.
That work has meant engaging with the spec as more than a schema to populate:
- We run LLM-based PDF-to-USDM conversion in production, which forces us to understand not just what each field is, but how the elements of a protocol actually relate to one another.
- We ported the CDISC CORE rules engine from Python to TypeScript and integrated validation directly into our application. That gave us granular familiarity with the conformance rules, where they are tightly specified and where edge cases surface.
- We have had to make deliberate decisions about when to extend the model for domain-specific data and when to fit within existing structure, a distinction that matters for interoperability across the DDF ecosystem.
Approaching the standard this way surfaces questions we think are worth bringing to governance. Reconstructing USDM from an unstructured document is inherently lossy. Some fields cannot be rebuilt from content alone and require interpretive reasoning. That raises a real design question for the spec: how should it represent provenance, confidence, and the rationale behind an inferred value, so that a downstream consumer can tell authored intent apart from reconstructed data?
It is also why we keep coming back to the gap between valid USDM and useful USDM. A study definition can pass every conformance rule and still be hard for the next system to actually use. Closing that gap is, to us, one of the more interesting problems in front of the standard. It is the same engineering effort that earned us a win in the TransCelerate Protocol Review Challenge.
What we hope to contribute
Bill is joining the UGG to bring that production, implementation-first perspective to the group: the view from teams actually building on USDM and running into its edges. Good standards are shaped as much by the people implementing them as by the people who design them, and we want to be useful on that side of the work.
Join us in Denver in October
If you will be at the CDISC US Interchange in Denver this October, you can hear more of this thinking in person. Bill is presenting "What It Actually Takes: Lessons from Production USDM Adoption and the Path to Digital Data Flow" in Session 6B on Tuesday, October 6. Come find us. And if you are working on USDM or Digital Data Flow, whether as a sponsor, a vendor, or a fellow implementer, we would like to compare notes.