The rule. If open ended answers matter to your remote study, build chatbot detection into the instrument before fielding. Screening is a design requirement now, not a cleanup step.
Why. In a survey of roughly 800 Prolific participants, 34 percent reported using LLMs to help answer open ended questions, and LLM text runs more homogeneous and more positive than human writing (Zhang, Xu, & Alvero, 2024). Traylor (2025) compared a managed panel to an MTurk sample and found the MTurk responses looked higher quality yet were more likely to be chatbot generated, even after screening on closed ended questions and paradata. Asher et al. (2026) embedded a keystroke logger in three Prolific studies (N = 928): about 9 percent of participants pasted or barely typed their answers despite deterrence measures, enough in simulation to shrink observed effect sizes by 10 percent and inflate required sample sizes by up to 30 percent.
Seen in the wild. Prolific's researcher help center now ships behavioral LLM checks (copy paste and tab switching signals, 98.7 percent precision in testing) and files traditional attention checks under what does not work (verified as of September 1, 2026).