Skip to main content
Storia TechnologiesResearch review · No. 012026 · Montréal, Canada

Building the knowledge graph layer for work that has to be defensible.

Most enterprise records are free text written about real things, assets, obligations, decisions, and then stored apart from the systems that hold those things. Our research programme is about closing that gap: linking unstructured language to the entity it is actually about, in a way you can audit.

The methods are domain-independent. They assume only that records are written in natural language about entities held somewhere else. We prove them on the hardest domain we can get real data for, then carry them outward.

Research tracks

3 method tracks. 1 domain, for now.

Updated as records are published
Method track

Knowledge representation

What the nodes and edges should be if a graph is going to stand for real-world entities, assets, obligations, events, rather than documents about them.

Foundational · no records yet
Method track

Semantic linking

Mapping unstructured language onto the specific entity it refers to. This is where our first published record sits, and the hardest part to do reliably.

1 record published
Method track

Interpretable AI

Keeping the language model out of the final decision. If a result cannot be explained step by step, it cannot be defended by an expert.

1 record published
Domain track

Construction & AEC

Our first vertical and current proving ground. It supplies the hardest datasets, RFIs, drawings, BIM models, and the clearest cost of a wrong answer.

1 record published
Domains where the methods run
Construction & AECInfrastructure & energyLegal & contract recordsManufacturing & asset ops

Only the first is active today: every published record below comes from construction. The others are stated as thesis, not as work in progress. The same problem shape, unstructured correspondence describing regulated physical or contractual entities, recurs in each, and the method tracks above are written to be domain-independent for exactly that reason.

Lead paper · EC3 2026

Semantic Mapping of Request for Information (RFI) to BIM Elements with Large Language Models

RFIs cause a lot of headaches. They're written in ordinary language, they carry no structure, and despite years of tooling around them the process is still slow, still manual, and still open to being gamed. This study looks at whether LLMs can help by automatically matching RFIs to elements of the BIM model. The approach keeps the language model out of the decision: it reads the RFI for intent cues, then a five-stage pipeline of identifier, spatial, and system checks ranks the candidate elements, and a person confirms the link. Where the prototype failed, it failed on metadata; missing room associations, absent identifiers, thin model detail. That's a data problem, and it's the one worth solving first.

Fig. 1The language model extracts intent cues and never selects an element. Five ordered stages, each a rule against BIM metadata, do the choosing; a person confirms the link. Adapted from the paper's prototype architecture.
Records describe real-world entities. The data does not connect them to those entities.

The gap this programme is aimed at. In construction it is an RFI and the BIM element it refers to, the case our first published record closes. The shape recurs wherever language describes something regulated.

The research-to-practice loop

01

Problems come from the field

We bring questions that actually block working teams, not synthetic benchmarks. The gap our first record addresses surfaced from real correspondence, not from a literature review.

02

Method comes from the university

Framing, validation design and peer review sit with the research group. That is the part a vendor cannot mark its own homework on.

03

Findings generalise past the domain

A method proven on the messiest domain we can find transfers outward more readily than one built on a clean one. The first domain is a stress test, not a ceiling.

Who we work with

All research

1 published
EC3 2026

Semantic Mapping of Request for Information (RFI) to BIM Elements with Large Language Models

RFIs cause a lot of headaches. They're written in ordinary language, they carry no structure, and despite years of tooling around them the process is still slow, still manual, and still open to being gamed. This study looks at whether LLMs can help by automatically matching RFIs to elements of the BIM model. The approach keeps the language model out of the decision: it reads the RFI for intent cues, then a five-stage pipeline of identifier, spatial, and system checks ranks the candidate elements, and a person confirms the link. Where the prototype failed, it failed on metadata; missing room associations, absent identifiers, thin model detail. That's a data problem, and it's the one worth solving first.

Semantic linkingInterpretable AIConstruction & AEC
Presented

Researching something adjacent? We'd like to hear about it.

We're open to collaborations with groups working on knowledge graphs, semantic retrieval, and evidence you can defend, in any domain where unstructured records have to be tied back to the entities they govern. We can bring real problem statements and anonymised datasets.

Contact the research team
Research: knowledge representation and semantic linking · Storia