What we ran · what held up · what we would change

Methods and lessons

The behind-the-scenes account for analysts: what we ran, what held up, and what we would do differently. Any unfamiliar term is in the plain-English glossary.


What we actually did

Study 1 is a two-wave critical-incident panel. Time 1 narratives (March 2016) were dual-coded on 13 Rosso facets. Time 2 (February–March 2017) measured felt meaningfulness, emotional exhaustion, and turnover intentions, the three outcomes the registered tests used, alongside other scales. The interval is about one year (M = 11.58 months). Completers were older and higher in work centrality than noncompleters.

Before any outcome model, we tested whether a language-model panel could replace the human pair. Model instructions were written only on the practice slice, locked (hash-recorded), and scored once on the try-out slice. No set-aside case was scored while any set of instructions could still be chosen. None reached the pass mark. Predictors stayed human. The one later test look is described below.

Studies 2 and 3 use Internet Archive snapshots of allnurses.com, captures January 2016 through 8 September 2026: 7,709 posts, 1,453 threads, 4,443 accounts. The instrument is the Study 1 codebook two humans built on prompted incidents. One model applied it to every post for meaning events (facet × enacted/frustrated/wished × object named × task fate × emotion × attribution) and four post-level flags; 5,435 events, two-thirds of posts with none. A 200-post transfer sheet for a second human coding of that forum prose was drawn and, by PI decision, not coded for this packet. Calls ran at the provider’s default temperature (the lock in the protocol was not applied). Dictionaries cover the full window as archive description. Intervals cluster on thread; account-level resampling was the same or tighter.

Ethics and provenance. Study 1 ran under expedited review by the University of San Francisco IRBPHS (#611); the Time 2 wave was authorized by a modification signed December 27, 2016. The consent form did not address quotation, so no verbatim narrative leaves the private repository; exemplars are paraphrased. Studies 2 and 3 analyzed publicly available forum pages compiled from Internet Archive snapshots: no contact with the site’s servers, no interaction with posters, no private identifiable information, author identifiers hashed at ingest, no post text reproduced. Under 45 CFR 46.102(e) that activity is not human-subjects research and was not submitted for IRB review (PI determination, September 25, 2026). Forum text is handled like respondent microdata: scripts read it, people read aggregates.

What worked

What we would do differently

Lessons we would hand to another lab

On measurement

Do not set one acceptance number for “meaning.” Agency and belonging are not the same coding problem. If humans cannot see a dimension, a model fitted to those humans will inherit the blindness and look confident.

On models

A published paper that says models beat crowd workers is not a validation of your codebook. The bar is your human pair, on your facets, on your text type. Forum text is a new validation, not a free extra N.

On outcomes

Concurrent convergence (codes with WAMI in the same wave) is not prospective prediction. We had the first and not the second. Report both or you will overclaim the map.

On time

Capture year is not posted date. Posted date is not AI exposure. Hospital adoption is not this writer. A 2023 cut is the wrong era start; the apparatus talk was there in 2016 at the same rate. Nurses do not say “AI”; they say ratios, metrics, charting. Measure the artefact, not the word. None of this is a causal AI test.

The independent panel (Session 6, Tier 3)

This was specified on 11 June 2026, not invented after the try-out run missed the pass marks. Tier 1 is cells both humans agreed on. Tier 2 is the blinded sample of disagreements. Tier 3 is the panel scored against Time 1 WAMI and Time 2 strain and retention — a criterion no coder, human or model, has seen. Human agreement on communion is .48. A model chosen to match that pair will carry the disagreement forward. A panel that is not selected on human agreement, and is judged on later outcomes, is a different test.

It was not run because ADR 0006 parked all further LLM scoring of the split after every set of instructions missed the human-agreement pass marks and the manuscript had a deadline. The park was a shipping decision. The June design still stands. ADR 0010 puts Tier 3 back on the table. After the manuscript results were fixed, one descriptive look at 65 set-aside cases was taken under a protocol written and hash-recorded beforehand (ADR 0012), on PI decision with coauthor deferral; it scored the two best of the rejected setups and is reported as a 65-case number at try-out precision, not as validation: agency .81 / .80, others-directed .746 / .770, still short of the human bar, salience reliability in the drop band. Same-case human-pair agreement on those 65 was .95 / .84, so the slice was not unusually hard. The pre-set rule for a later single look at the remaining 128 cases (agency and others-directed both ≥ .75) missed on others-directed at .746; the 128 remain sealed and refused by code. Train was the prompt sandbox. Dev is spent. Twenty-nine narratives sit outside the split (eight with Time 2) and can be scored without an unseal; that is a mechanics check, not the powered test.

No respondent text or forum post text is in this packet. The 128-case test remainder stays sealed. If you want the term definitions, use Plain-English glossary.