Color digital poster · one scrolling page · SIOP LEC 2026

The AI Paradox

Organizations automate helping work to free people for meaningful care. That only works if someone already knows which tasks carry meaning. We mapped the meaning, then tested whether a language model could take over the mapping. The map was thinner than the efficiency story. The model did not qualify. Then we read a decade of nurses’ own talk: they never said “AI.” They named ratios, metrics, and charting, and what rose in that talk was a sense of lacking control over means.


Matthew J. Monnot1 · Isaac Thompson2 · 1Vocentive · 2Amazon

One sentence

If optimization removes the tasks through which people experience meaning, efficiency is paid for in strain and exit.

The argument

Meaning is task-bound

Rosso, Dekas, and Wrzesniewski (2010) organize meaning as Agency–Communion crossed with Self–Others. Contribution and unification are two of those pathways. They are also the tasks now first in line for documentation, monitoring, and coordination tools.

More meaning is not always better

Calling-infused work can bind people to exhaustion (Bunderson & Thompson, 2009). Rosso and colleagues left open a curvilinear conjecture: too many active pathways may exhaust rather than sustain. A full reversal into strain is almost never identified. What this sample shows is flattening: Time 1 felt meaning rises from zero to two pathways and then levels.

Meaning can be measured at scale once the ceiling is known

People analytics will be asked to score meaning from text, and this study supplies the working method: find out how well two trained humans agree on each kind of meaning first, write the pass marks down and set aside a file of cases before any model is chosen, and accept a model only where it clears the mark. Agency and others-directed meaning already have ceilings high enough to test against; communion needs better human agreement before any model can be judged on it. That map is what turns text scoring at scale from a vendor claim into an auditable measurement.

The 2 × 2 we coded

Thirteen present/absent facets nest in four pathways. A pathway counts as present if any of its facets is. We do not score Rosso’s seven mechanisms as a middle layer; they cut across the 2 × 2. Color on the figure is the pathway: navy agency, teal communion, grey Self–Others.

2 by 2 meaning pathways with facet lists
Board figure: the same 2 × 2 as a single graphic.

The design

Time 1 (March 2016): critical-incident narratives and meaning surveys, helping-profession workers, healthcare-majority online panel, N ≈ 307. Time 2 (February–March 2017): felt meaningfulness, emotional exhaustion, and turnover intentions, n = 144 complete longitudinal cases. Two trained humans coded 276 stories on all 13 facets (3,614 cells). Disagreements went to blinded adjudication. Model instructions were written only on a practice slice of the stories and scored once on a try-out slice. After the paper’s results were fixed, one descriptive look at 65 of the set-aside cases was taken under rules written down beforehand (two setups that had already missed the mark; reported as a number on cases the models had never seen, not as validation); the remaining 128 cases stay untouched.

Four-step study design timeline
How the stories were split: practice ~10% / try-out, looked at once ~20% / set aside ~70%.

Finding 1. The human ceiling is not one number

How often the two coders agreed a kind of meaning was present when either saw it, with 95% intervals; n = 276 stories. Position on 0–1, not bar length.

Agency
.93 [.91, .96]
Others
.84 [.80, .88]
Communion
.48 [.34, .60]
Self
.52 [.43, .59]

Vendor claim: “we measure belonging from text at .90.” The human pair in this codebook does not reach that on communion. Ask which ceiling the product beat, on which dimension. Color is the pathway: navy agency, teal communion, grey Self–Others. Communion’s interval is the mark that matters.

Finding 2. None of six locked setups reached the agency bar

None of six locked setups reached the agency bar. Agency .750–.798 against .910. Highest .798 is .11 below the bar.

Six frozen prompts: agency and others-directed agreement against the two locked acceptance bars
Finding 2. None of six locked setups reached the agency bar. Filled dots: agency; open dots: others-directed. Vertical lines: pass marks (.910, .805). Agency .750–.798; highest .798 is .11 below .910.

A model chosen because it matches one human pair inherits that pair’s disagreements. The next test, designed in June, judges a panel of models against later well-being and retention rather than against human agreement (ADR 0010); it has not touched the 128 set-aside stories.

Finding 3. Meaning narrated as competence reached next year’s strain

When Time 1 stories grounded meaning in competence or mastery, Time 2 turnover intentions were higher (b = 1.40, 95% CI [0.58, 2.22], p < .001, n = 100) and so was emotional exhaustion (b = 0.95 [0.20, 1.70], p = .014, n = 94). Seventeen of 140 matched cases carried the code, and the association held after adjusting for the full Time 1 survey set, which itself forecast Time 2 meaning strongly. The planned next-year paths from the three broad meaning dimensions came in near zero (three of five inside the band we treat as no effect; Table 2 in the manuscript), which is what singles out competence as the signal that reached next year. It is the lead for the next designed study and should be read as a lead, not a rule.

Seventeen of 140 matched respondents whose Time 1 story grounded meaning in competence, with Time 2 turnover intention and exhaustion group means and 95 percent intervals
Finding 3. One mark per matched respondent (17 of 140 present-coded; counts only, no scores are drawn). Group means with 95% CIs for Time 2 turnover intentions and emotional exhaustion; adjusted estimates are in Table 4 of the manuscript.

Concurrently, the same codes track self-reported meaningfulness: others-directed r = .24, agency r = .20, contribution r = .23. Felt meaning rises with the number of pathways a person narrates up to two, then levels.

Pathway density and concurrent WAMI
Time 1. The more pathways a story carries, the higher the same person’s survey meaning score, up to two; then it levels.

Finding 4. Nurses named the apparatus, not AI

Internet Archive snapshots of allnurses.com, 7,709 posts of at least 200 characters in 1,453 threads, capture dates 1 January 2016 to 8 September 2026 (2026 is a partial year, 147 posts). The codebook is the one two trained humans built on the critical-incident stories. One language model applied that instrument to each post for meaning events: a Rosso facet enacted, frustrated, or wished for, plus what the post blamed or credited (a named AI system, a system rollout, or a management practice such as ratios, productivity metrics, dashboards, or documentation load) and four post-level flags (quit intent, lack of control, anxious, depressed). A 200-post transfer sheet for a second human coding of that forum prose is drawn and not yet coded. The year is when the archive copy was taken.

Seven events in 7,709 posts named an AI system; none after 2021. About one post in twenty named the optimization apparatus, and that share did not move across the decade (4.8% → 5.2%, 95% CI on the difference [−1.8, 2.4]). The phenomenon the AI-paradox literature describes was in nurses’ talk the whole time, under management’s vocabulary. Measuring “AI exposure” by whether workers say AI would miss it; the exposure that matters is the metrics-and-scheduling apparatus that AI now powers.

Object talk, frustrated events, and control-loss flags by capture year, 2016 through September 2026
Finding 4. Share of posts by archive year, with 95% intervals. Talk about ratios, metrics, and charting is flat; named AI near zero. The shaded band is a 2019–2021 site redesign that shortened posts; every line is read with that caveat.

Finding 5. What moved was control, not skill

Half of the posts naming ratios, metrics, or charting landed on control over how the work is done; frustrated control was the only category that grew (9.6 → 12.7 events per 100 posts; p = .028, but not once we correct for testing 20 categories). Posts flagged for lack of control rose from 10.4% to 14.5% (+4.0 points [1.2, 6.8]; unchanged after adjusting for post length; the only one of eight post-level shares that holds up after correction, q = .028). Competence stayed the most often expressed source of meaning in every period, and belonging and connection were expressed far more often than frustrated. Within the badged posts, RN exposure was flat and advanced-practice exposure rose with an interval that includes zero.

Object exposure by career-point badge and capture period
Finding 5. Role is read from the author-pane credential badge at capture, never from post text; 48% of posts carried a badge. Cells under 40 posts are marked.

Posts that name ratios, metrics, or charting also carry quit and lack-of-control flags more often (same post, same model). The language of belonging runs through nearly every post that carries any event yet rarely shows up as a separate episode; we read it as the standard nurses judge lost control against, a hypothesis for a designed study.

What to do this quarter

  1. Set aside a file of comments nobody scores until the model is final.
  2. Write the pass mark down first, one per kind of meaning, from how well two trained humans agree. Do not accept one overall accuracy number.
  3. Make the scoring job check itself: no set-aside comments in the practice pile, no names or identifying details in the output.

The longer kit is in Try this quarter.

Print tiles: hang slides 1–16. Slides 17–18 are extra letter pages (survey meaning by number of pathways; forum categories). Open Monnot_Poster_v5.pptx.

Study 1 maps meaning; AI exposure is read from the documented automation frontier, not measured per person. Forum results are single-model event codes, descriptive.