Two shapes for the value set — relation or function — and a plan to branch the reference implementation — Sep 15, 2026
Two shapes for the value set: relation or function
Nikolai Ryzhikov came in with a written follow-up to last week's terminology debate — this time with two concrete candidates on the table.
Candidate one — value set as a relation. A pre-expanded value set loaded into the database as a flat table, referenced via relatedArtifact inside a ViewDefinition. The runner injects a common table expression backed by a shared value-set table with the minimal columns — system, version, code, display, inactive. Queries then treat it as any other table: join, count, intersect, introspect. Optional properties can be added as extra columns when needed.
Candidate two — value set as a function. A memberOf(<url>) function inside FHIRPath that hides the mechanism. Under the hood it can be optimised with bitmaps, Bloom filters, or backed by a materialised expansion of the same shape — but the surface is a function call, not a table.
Nikolai's own summary of the trade-off: relations are transparent and cheap to implement (call terminology server, load the expansion, done), while functions can be much more performant with the right internal optimisation but hide the mechanism. Neither is obviously superior; both can coexist.
Somebody likes functions, somebody likes relations. Maybe we need both, or maybe we don't. It's not kind of interoperable — we have to recommend one of them, maybe support both, or throw one away.
Gino Canessa raised the FHIRPath alignment question: Brian Postlethwaite's proposed cleanup uses a JSON parameters blob rather than discrete arguments to accommodate exotic terminology inputs. If the SQL side goes with a memberOf function, matching signatures would help portability — but Gino worried about the combinatorial explosion of parameterisation across a dozen property flags.
Where relations pull ahead: non-selectable codes and extended properties
Gino pointed the group at a live Zulip argument about a non-selectable ICD-10 code that John Murkey was trying to use — and admitted he was horrified to find the spec says "not selectable" without documenting what that actually looks like in practice.
The relational shape sidesteps this cleanly. If you care about non-selectable, inactive, deprecated, or any other code-system property, it becomes a column at extract time and filtering happens in normal SQL:
The functions don't allow that as nicely. If we think we're going to want properties like that, I would lean towards a with-value-set-as-a-relation — because then, if you want to add different property values and filters, that's very easy.
Gino's follow-on was that a ViewDefinition-shaped template could formalise the basic value-set table (system, version, code, display, inactive). Anything more exotic — an extra property, a filter, a diabetes-codes-with-severity variant — becomes a specific ViewDefinition on top. Simple template for the common case, escape hatch for the tail.
RxNorm, SNOMED, and hiding the distribution nightmare
Gino brought up RxNorm's database scripts as a data point on the long tail — Oracle and MySQL tables with loading scripts you can just download without login. John Grimes noted SNOMED CT ships in the same shape: a tabular distribution plus, in his phrase, "an absolute nightmare of SQL" you're supposed to run over it to make it queryable.
That, John argued, is precisely the problem the abstraction solves. Ontoserver's whole value proposition was to stop people running SNOMED prep SQL over and over. A SQL on FHIR runner does the same job for analytics: whatever the source terminology looks like (RxNorm, SNOMED, LOINC), a loader translates it into the shared table-based interface — and above that line, the query author never has to know.
It's a terrible idea to give the average analytic user those RxNorm SQL files and say this is what you have to do.
Nikolai's follow-on: the deeper relational terminology model (concept, relationship, property, description, designation) can express RxNorm, SNOMED, and most of FHIR terminology — Termbox uses this and covers around 90% of real-world terminology needs. But that is the next-level problem. The value-set abstraction being discussed today sits on top of it and does not require solving distribution first.
John's read on Athena is that the ViewDefinition-and-value-set framing is actually cleaner than what OHDSI ended up with. Athena distributes concepts, relationships, and properties as tables, but the value-set materialisation happens later in Atlas via a separate library that generates complex SQL at runtime. Anchoring on the value set instead — the way this proposal does — means the composition logic can hide behind the abstraction rather than being smeared across the query author's code.
Version resolution: the fuzzy default and its cost
Owen Loveluck raised a practical concern: both proposals reject ambiguous versioning, but real value sets often ship without pinned versions. If two versions of the same value set land in the database, what does the join return?
Gino's answer was blunt: use explicit versions. If your authors aren't pinning, go back to your authors. Owen's counter was equally blunt — he already has the un-pinned value sets and the publishers don't care.
Nikolai's team learned the lesson in production. Aidbox's validator once fell back to the "latest" resolution for encounter-status codes; a breaking change to the code system silently broke validation for every R4 server they ran. They pin explicitly everywhere now.
But how do you know what latest is? If there's no version, you need some way of ordering versions.
John's larger point: different code systems have different versioning schemes and different semantics. SNOMED CT has concept permanence, so unpinned is actually safer than pinning. Others break silently across releases. There is no single ordering algorithm that works everywhere — the R5+ versionAlgorithm extension exists precisely because you have to know each code system's convention (SemVer, date, alphanumeric, DICOM's year-plus-letter, etc.) to order versions at all.
Branching the reference implementation, and drawing the scope line
Steve Munini's preference was on the record early: relations feel more straightforward, but the group can implement both. John moved that the way to settle it is to branch the reference implementation and actually build something.
Two open scope questions John flagged before committing:
- Should
translate,memberOf, and other terminology operations be part of ViewDefinition itself, or an optional runner capability? - What is the full scope of the SQL-side terminology abstraction — just membership, or does it also cover subsumption, mapping, properties, and designations?
Nikolai's take: value-set membership is the primary case, so start there and let the scope line emerge from what implementations actually need. John agreed — the bones of the proposal look right, and keeping the value-set composition (turning a compose into SQL) behind the abstraction rather than pushed into query authors' laps is what he thinks brings Grahame Grieve along.
ViewDefinition resource proposal — FMG approval landed
Gino confirmed the ViewDefinition resource proposal was approved by FireEye. He was going to bring it to FMG this week, but the WGM collapsed the agenda; it will come up either at the working group meeting or immediately after. The only complaint from the reviewers, delivered with a straight face: John's documentation is so succinct yet thorough that everyone else's submissions now look bad by comparison.
John cannot make the WGM in person; the group agreed to keep the following meeting on the calendar anyway and use the interval to shape the proposal into a branch.