If you are about to score employee text with a model
- Set aside a file of comments nobody scores until the model is final. Ten percent, locked, never used as examples. Every look at it spends it. Validation done on comments that were also used to improve the instructions reports a precision that does not exist (Törnberg, 2024); this is the same discipline selection researchers use when they hold out test data and swap it only once (Koenig et al., 2023).
- Write the pass mark down first, one per kind of meaning, from how well two trained people agree. In our data two coders agreed .93 on agency and .48 on belonging. Agency and belonging will not share a mark. If two trained people cannot see it, do not buy a model that claims it.
- Keep humans as gold until a model clears the mark they set. No set-aside comments in the practice pile, and no names or identifying details in the output. None of six locked setups reached the agency bar. The paper ran on the human codes.
If you are about to automate part of a helping job
- Map the meaning before you pick the tool. List the tasks the tool will touch and, for each, ask whose meaning it carries and which of three things the tool does to it: replaces it, leaves a person tending the machine, or amplifies what the person can do (Bankins & Formosa, 2023). Work-design choices made at implementation, not after, are what decide whether autonomy, skill use, and feedback survive the rollout (Parker & Grote, 2022). The worksheet below is for this step.
- Keep the decision with the person, and make the tool explainable to whoever follows it. In nurses’ own talk, what rose over a decade was not skill loss but loss of control over how the work is done. Critical-care staff asked for the same thing when interviewed about AI: keep decision-making local and keep the system transparent and predictable (Bienefeld et al., 2025). Recommendations workers cannot see the basis of are complied with symbolically, not followed (Kellogg et al., 2020); people should be held accountable only for processes they can understand, predict, and influence (Parker & Grote, 2022).
- Pilot small, with the workers, and measure meaning, not just minutes. The ambient-scribe site that reported early time savings ran a two-week pilot with 47 clinicians before a regional rollout that named physician champions and a feedback loop (Tierney et al., 2024). The site that reported burnout measured it before and after, not just documentation time (Shah et al., 2025). Add two items to your pre/post survey: control over how I do my work, and chance to use my skills. If those fall while minutes improve, you have found the paradox in your own building.
Twelve weeks
| Weeks | Text scoring track | Automation track |
|---|---|---|
| 1–2 | Pull the comment set. Set aside 10% in a locked file with a dated note of who may open it and when. Decide which kinds of meaning you actually need scored. | List the six tasks the tool will touch. Fill in the meaning map with the team that does the work, not for them. |
| 3–4 | Two trained people code 100 comments from the practice pile, per kind of meaning. Compute agreement per kind. Write the pass marks down and date them. | For each task, decide: replace, tend, or amplify. Name the decisions that stay with the person. Draft the two pre/post survey items. |
| 5–8 | Build or buy. Tune only on the practice pile. Log every run: model version, instructions, date. | Two-week pilot with volunteers and one named champion per unit. Pre-survey first. Collect what people work around, not just what they praise. |
| 9–10 | One look at the try-out cases against the written pass marks. No edits to the instructions after seeing the numbers. | Post-survey. Compare minutes saved against control and skill-use items. Review what the champions heard. |
| 11–12 | Decide: accept where the mark is cleared, keep humans where it is not. Write the one-page record. The set-aside file stays shut unless you accepted. | Decide: scale, redesign, or stop, task by task. Write down which decisions stayed with the person and why. |
Ask the vendor
- Which two humans did the model get compared to, and on which kinds of meaning?
- When the model says a kind of meaning is present, how often did the humans agree? Overall accuracy hides this, because most comments carry nothing.
- How often did the model find meaning in comments both humans said had none?
- Was the check done after the instructions were final, on comments that were never used to write them? How many comments?
- Could the model have seen the check data before? Public datasets do not count (Törnberg, 2024).
- Does the product report belonging or authenticity? How well did two humans agree on those?
- If the model is wrong, how does that error enter the dashboard that managers see?
If they answer with one accuracy number and a crowd-worker paper, send them Gilardi et al. (2023) and Törnberg (2024) together. The first is their citation. The second is the standard.
Ask the tool vendor
- Which decisions does the tool make, which does it recommend, and which does it leave alone? Show me the list.
- Can the person using it see why it recommended what it did, in their own terms?
- What does it record about the worker, who sees that, and is any of it used for performance review?
- Which sites piloted it, for how long, and what did they measure besides time saved?
- What can the worker override, and does overriding get logged against them?
A 20-minute meaning map
In a team that is about to adopt a documentation or monitoring tool, list six tasks. For each, mark: whose meaning does it carry (self, other, both, unclear)? What does the tool do to it (replace, tend, amplify)? If it replaces the task, what occasion for meaning disappears, and what is the replacement? This is a discussion device, not a validated instrument. In our data the broad meaning dimensions did not forecast next-year strain; stories anchored in competence did, in 17 cases, and that lead is exploratory. Treat the column “what is lost” as the hypothesis you will test in the pilot survey.
| Task | Meaning source (if any) | Replace / tend / amplify | Decision stays with the person? | What is lost if it goes? |
|---|---|---|---|---|
Print this page. The empty rows are the point. Plain definitions of the set-aside file and the agreement score sit in the Plain-English glossary.
Where this comes from
All of these are in the project’s literature folder and were read for this page. The two trade pieces on ambient scribes are implementation reports, not trials; take their numbers as early experience.
- Bankins, S., & Formosa, P. (2023). The ethical implications of artificial intelligence (AI) for meaningful work. Journal of Business Ethics, 185(4), 725–740. https://doi.org/10.1007/s10551-023-05339-7
- Bienefeld, N., Keller, E., & Grote, G. (2025). AI interventions to alleviate healthcare shortages and enhance work conditions in critical care: Qualitative analysis. Journal of Medical Internet Research, 27, e50852. https://doi.org/10.2196/50852
- Gilardi, F., Alizadeh, M., & Kubli, M. (2023). ChatGPT outperforms crowd workers for text-annotation tasks. Proceedings of the National Academy of Sciences, 120(30), Article e2305016120. https://doi.org/10.1073/pnas.2305016120
- Kellogg, K. C., Valentine, M. A., & Christin, A. (2020). Algorithms at work: The new contested terrain of control. Academy of Management Annals, 14(1), 366–410. https://doi.org/10.5465/annals.2018.0174
- Koenig, N., Tonidandel, S., Thompson, I., Albritton, B., Koohifar, F., Yankov, G., . . . Newton, C. (2023). Improving measurement and prediction in personnel selection through the application of machine learning. Personnel Psychology, 76(4), 1061–1123. https://doi.org/10.1111/peps.12608
- Parker, S. K., & Grote, G. (2022). Automation, algorithms, and beyond: Why work design matters more than ever in a digital world. Applied Psychology, 71(4), 1171–1204. https://doi.org/10.1111/apps.12241
- Shah, S. J., Devon-Sand, A., Ma, S. P., Jeong, Y., Crowell, T., Smith, M., Liang, A. S., Delahaie, C., Hsia, C., Shanafelt, T., Pfeffer, M. A., Sharp, C., Lin, S., & Garcia, P. (2025). Ambient artificial intelligence scribes: Physician burnout and perspectives on usability and documentation burden. Journal of the American Medical Informatics Association, 32(2), 375–380. https://doi.org/10.1093/jamia/ocae295
- Tierney, A. A., Gayre, G., Hoberman, B., Mattern, B., Ballesca, M., Kipnis, P., Liu, V., & Lee, K. (2024). Ambient artificial intelligence scribes to alleviate the burden of clinical documentation. NEJM Catalyst Innovations in Care Delivery, 5(3). https://doi.org/10.1056/CAT.23.0404
- Törnberg, P. (2024). Best practices for text annotation with large language models (arXiv:2402.05129). arXiv. https://doi.org/10.48550/arXiv.2402.05129