Unitree robot inspecting rust in a subway tunnel
AI image (openAI): Unitree robot inspecting rust in a subway tunnel

The Subjectivity of Severity: Can We Trust LLMs to Label Infrastructure Damage?

Trai­ning a super­vi­sed ML model requi­res label­led data. Label­led data requi­res domain experts. In spe­ci­fic domains like infra­struc­tu­ral inspec­tion, the­se experts are scar­ce, expen­si­ve and pro­ba­b­ly have bet­ter things to do than cli­cking through hundreds of images.

So, I asked a logi­cal ques­ti­on: Could we make use of LLMs to eva­lua­te rust? And would they per­form bet­ter than an abso­lu­te ama­teur like me? I deci­ded to give it a try and star­ted an expe­ri­ment with seve­ral Visu­al Lan­guage Models (VLMs), Lar­ge Lan­guage Models with visi­on enco­der fea­tures.

The set­up: 291 rust images from the DACL10k vali­da­ti­on data­set, rated on a con­ti­nuous 0–1 seve­ri­ty sca­le. Four raters: me (data sci­en­tist, zero inspec­tion expe­ri­ence), GPT-4o, Clau­de Son­net, and GLM-4.6V-Flash-9B run­ning local­ly.

All four recei­ved the exact same prompt: “Act as an expe­ri­en­ced struc­tu­ral inspec­tor, rate rust seve­ri­ty accor­ding to DIN 1076 (the Ger­man stan­dard for inspec­ting engi­nee­ring struc­tures, rating dama­ge by its struc­tu­ral signi­fi­can­ce and urgen­cy of action.), use the full 0–1 ran­ge across six anchor cate­go­ries from insi­gni­fi­cant sur­face dis­co­lou­ra­ti­on to seve­re sec­tion loss, and return struc­tu­red JSON.”

Com­pa­ring the out­puts:

The Subjectivity of Severity: Can We Trust LLMs to Label Infrastructure Damage?
Figu­re 1. Pair­wi­se bias and cor­re­la­ti­on bet­ween four raters (Self, GPT-4o, Clau­de Son­net, GLM-4.6V) on 291 rust seve­ri­ty labels (0–1 sca­le).

The seve­ri­ty orde­ring: Ope­nAI < Self < Clau­de ≲ GLM. GPT-4o is the most leni­ent while Clau­de and GLM share near­ly iden­ti­cal mean ratings (bias = −0.009) despi­te only modera­te­ly agre­e­ing on indi­vi­du­al images (r = 0.60, MAE = 0.133). I sit some­whe­re in the midd­le (which I choo­se to inter­pret as being well-cali­bra­ted and not as being inde­cisi­ve).

Ope­nAI and Clau­de have the stron­gest rank cor­re­la­ti­on (r = 0.77), they most­ly agree on which images are more seve­re than others, but Clau­de rates about 0.18 hig­her on avera­ge. One thing that stands out is that GLM, an open-source model, ending up in the same seve­ri­ty are­as as Clau­de is both sur­pri­sing and rele­vant if you’­re working with data that can’t lea­ve your net­work.

The cor­re­la­ti­ons across all pairs ran­ge from 0.40 to 0.77, show­ing some shared signal and not just ran­dom noi­se. But the­re’s also real dis­agree­ment in both avera­ge seve­ri­ty level and in how indi­vi­du­al images get ran­ked. If raters only agree to this ext­ent, that agree­ment sets a cei­ling on what any down­stream model can learn from the­se labels.

The Subjectivity of Severity: Can We Trust LLMs to Label Infrastructure Damage?
Figu­re 2. Seve­ri­ty rating com­pa­ri­son examp­les

Should we let LLMs label our ground truth? 

If I have zero pro­fes­sio­nal expe­ri­ence in struc­tu­ral engi­nee­ring, why shouldn’t we just accept Clau­de’s or OpenAI’s ratings as the abso­lu­te ground truth?  

The short ans­wer is: Becau­se “con­sis­t­ent­ly con­fi­dent” is not the same as “cor­rect.” 

A visu­al model doesn’t feel the humi­di­ty; it doesn’t under­stand the phy­si­cal con­text of the sur­roun­ding con­cre­te. It is trans­la­ting pixels into text based on pro­ba­bi­li­ty, not phy­sics. If we blind­ly accept VLM out­puts as ground truth, we risk trai­ning com­pu­ter visi­on models on “hal­lu­ci­n­a­ted urgen­cy” or “hal­lu­ci­n­a­ted safe­ty.” 

What I’d like to explo­re next:

  • Can an ite­ra­ti­ve approach help e.g. AI rates first, human reviews only the uncer­tain ones
  • Whe­ther the bia­ses are prompt-cor­rec­ta­ble or archi­tec­tu­ral
  • What inter-rater agree­ment looks like bet­ween actu­al human inspec­tors on the same images (could someone from the inspec­tion world vali­da­te this? Genui­ne ask.)

If cer­ti­fied inspec­tors also agree at r ≈ 0.7, then the label noi­se isn’t a data qua­li­ty pro­blem, but a pro­per­ty of the domain its­elf.

The reason I’m stuck label­ling rust images in the first place is Robo­TUNN, a pro­ject whe­re a robot dog inspects sub­way tun­nels and feeds the data into a digi­tal twin for pre­dic­ti­ve main­ten­an­ce. Befo­re any of that can work, someone has to deci­de what counts as seve­re dama­ge. Right now, that someone is me, three LLMs, and a lot of dis­agree­ment.

If you work in infra­struc­tu­re inspec­tion and have thoughts on this, or access to label­led dama­ge data, I’d love to talk.

Categories
Latest News
Your contact

If you have any ques­ti­ons about our pro­ducts or need advice, plea­se do not hesi­ta­te to cont­act us direct­ly.

Tel.: +49 40 22 85 900-0
E-mail: info@statsoft.de

Gui­do Band­holz (Head of Sales)