Bibliography
Crema's design decisions rest on published research in software inspection, sustained attention, language-model behavior, requirements volatility, and oversight. Each entry below gives the citation, the finding in one or two sentences, a quote where we have read the source, and the product decision it supports.
Every citation was checked against its DOI record on 14 September 2026. Entries marked VERIFIED were read at the source and carry a quotation.
-
VERIFIEDREVIEW RATE
Kemerer, C. F., & Paulk, M. C. (2009). The impact of design and code reviews on software quality: An empirical study based on PSP data. IEEE Transactions on Software Engineering, 35(4), 534–550.
TL;DRTwo data sets, 371 and 246 programs. Review rate predicted defect removal after controlling for developer ability. At 200 lines an hour or slower, reviews caught nearly two-thirds of design defects and over half of code defects.
“The recommended review rate of 200 LOC/hour or less was found to be an effective rate for individual reviews, identifying nearly two-thirds of the defects in design reviews and more than half of the defects in code reviews.”
D1 · The governed unit is the change, not the documentD4 · Mandates carry numeric limits -
VERIFIEDATTENTION
Pattyn, N., Neyt, X., Henderickx, D., & Soetens, E. (2008). Psychophysiological investigation of vigilance decrement: Boredom or cognitive fatigue? Physiology & Behavior, 93(1–2), 369–378.
TL;DRReaction times slowed after 30 minutes on task, with a further cost to self-directed (endogenous) attention after 60 minutes. Physiological measures pointed to underload, not exhaustion.
“The vigilance decrement has been described as a slowing in reaction times or an increase in error rates as an effect of time-on-task during tedious monitoring tasks.”
D15 · The review clock is a feature -
VERIFIEDATTENTION
McCarley, J. S., Gyles, S. P., Hankey, K. R., Muñoz Gomez Andrade, F., & Yamani, Y. (2026). Deconstructing the vigilance decrement: Changes in bias, lapse rate, and guess rate, but not sensitivity. Attention, Perception, & Psychophysics, 88, 140.
TL;DR366 participants, 40 trials a minute in noise. Over time, people shifted to a conservative criterion, lapsed more, and guessed positive less. They did not lose the ability to tell signal from noise. A tired reviewer sees as well and reports less.
“Model comparisons revealed decisive evidence for conservative bias shifts, increases in mental lapse rates, and decreases in positive guess rate over time, but gave evidence against a loss of sensitivity.”
D7 · Approve, deny, regenerate, edit; no silent editsD15 · The review clock is a feature -
VERIFIEDAI-WRITTEN TEXT
Zhang, Y., Das, S. S. S., & Zhang, R. (2024). Verbosity ≠ veracity: Demystify verbosity compensation behavior of large language models. arXiv.
TL;DRFive QA datasets, 14 models. When uncertain, models repeat the question, hedge, and over-enumerate; the behavior appeared in every model and dataset. GPT-4 showed the behavior 50.40% of the time, and the gap between verbose and concise answers held as model capability rose.
“When unsure about an answer, humans often respond with more words than necessary, hoping that part of the response will be correct. We observe a similar behavior in large language models.”
D5 · Length is a drift signal -
VERIFIEDAI-WRITTEN TEXT
Zhu, J., Peng, H., Wang, J., Ke, L., Zhang, C., & Zhang, L. (2026). Large language models do not always need readable language. arXiv.
TL;DRText written for a model reader kept 99.5% of its meaning at 27.9% of its length. Human readability and model recoverability are separate properties. One document cannot serve both readers well.
“maintaining 99.5% semantic fidelity even when the text volume is condensed to 27.9% of its original length.”
D2 · One source, two renderings -
VERIFIEDAGENT OUTPUT
Duma, K., Wróblewski, P., Bobińska, J., Winiarska, J., & Przymus, P. (2026). These aren't the reviews you're looking for: How humans review AI-generated pull requests. arXiv. EASE 2026.
TL;DRAIDev dataset, 33,596 agent-authored pull requests. Most received no review; when reviewed, other agents did most of it. In repositories with both kinds of PR, 69.9% of agent-authored PRs had no recorded review or agent-only review, and human-only review was rare (8.08%).
“the absence of recorded review activity does not imply the absence of human oversight; maintainers may inspect pull requests without leaving traceable comments.”
P · Problem statement -
VERIFIEDOVERSIGHT
Turan, E. (2026). Oversight has a capacity: Calibrating agent guards to a subjective, fatiguing human. arXiv.
TL;DR125 hand-labeled agent actions. Reviewers agreed only moderately on what was risky (Fleiss' κ = 0.52). With the reviewer modeled as fatiguing, realized safety is an inverted U in escalation rate: past the optimum, more approvals make a system less safe. The author states the fatigue curve is simulated, not measured.
“Agent oversight, framed this way, is not only a classification problem but a resource-allocation one: human attention is finite, and the guard's escalation policy spends it.”
D8 · Approvals are rationed and de-duplicatedD11 · Measured churn, never a self-rated risk score -
DOI CHECKEDREVIEW RATE
Vitharana, P. (2017). Defect propagation at the project-level: Results and a post-hoc analysis on inspection efficiency. Empirical Software Engineering.
TL;DRInspection efficiency rises to an optimum with time invested, then falls. Defects missed at the requirements stage propagate into design and code.
D1 · The governed unit is the change, not the document -
DOI CHECKEDREVIEW RATE
Laitenberger, O. (2001). Cost-effective detection of software defects through perspective-based inspections. Empirical Software Engineering.
TL;DRInspection cost-effectiveness falls on large documents, attributed to fatigue and lost motivation. Inspect by logical entity, not by document.
D1 · The governed unit is the change, not the document -
DOI CHECKEDATTENTION
Reteig, L. C., van den Brink, R. L., Prinssen, S., Cohen, M. X., & Slagter, H. A. (2018). Sustaining attention for a prolonged period of time increases temporal variability in cortical responses. bioRxiv.
TL;DRAn 80-minute sustained attention task. Performance and motivation fell fast and reached a floor well before 60 minutes. A monetary incentive at 60 minutes restored motivation but not performance.
D15 · The review clock is a feature -
DOI CHECKEDATTENTION
Robison, M. K., & Nguyen, B. (2023). Competition and reward structures nearly eliminate time-on-task performance decrements: Implications for theories of vigilance and mental effort. Journal of Experimental Psychology: Human Perception and Performance, 49(9), 1256–1270.
TL;DRCompetition and reward nearly eliminated time-on-task decrements, which argues the decrement is motivational rather than a depleted resource.
D15 · The review clock is a feature -
DOI CHECKEDAI-WRITTEN TEXT
Saito, K., Wachi, A., Wataoka, K., & Akimoto, Y. (2023). Verbosity bias in preference labeling by large language models. arXiv.
TL;DRGPT-4 as a judge prefers longer responses more than humans do. Using a model to review a model-written spec rewards the padding.
D5 · Length is a drift signal -
DOI CHECKEDSPEC SIZE
Loconsole, A. (2008). A correlational study on four measures of requirements volatility. Proceedings of the 12th International Conference on Evaluation and Assessment in Software Engineering (EASE 2008).
TL;DRThe size of use-case models predicted the number of changes to them. Longer specifications change more often.
D10 · Drift detection uses size-scaled thresholds -
DOI CHECKEDSPEC SIZE
Loconsole, A., & Börstler, J. (2005). An industrial case study on requirements volatility measures. APSEC 2005.
TL;DRStakeholders' sense of volatility did not track measured volatility. Measure it directly.
D11 · Measured churn, never a self-rated risk score -
DOI CHECKEDSPEC SIZE
Ahrens, M., Schneider, K., & Kiesling, S. (2016). How do we read specifications? Experiences from an eye tracking study. Requirements Engineering: Foundation for Software Quality (REFSQ 2016), Lecture Notes in Computer Science, 301–317.
TL;DREye-tracked readers of specifications read by role. A larger specification is not always a better one.
D2 · One source, two renderings -
DOI CHECKEDSPEC SIZE
Tu, Y.-C., Tempero, E., & Thomborson, C. (2016). An experiment on the impact of transparency on the effectiveness of requirements documents. Empirical Software Engineering.
TL;DRReaders of documents rated higher on accessibility, understandability, and relevance answered faster, more correctly, and with more confidence.
D2 · One source, two renderings -
DOI CHECKEDSPEC SIZE
Basili, V. R., Green, S., Laitenberger, O., Lanubile, F., Shull, F., Sørumgård, S., & Zelkowitz, M. V. (1996). The empirical investigation of perspective-based reading. Empirical Software Engineering, 1(2), 133–164.
TL;DRNASA teams reading a requirements document from assigned perspectives covered more of it than teams reading their usual way.
D2 · One source, two renderings -
DOI CHECKEDSPEC SIZE
Kulk, G. P., & Verhoef, C. (2008). Quantifying requirements volatility effects. Science of Computer Programming.
TL;DR84 projects. The healthy rate of requirements change depends on project size and duration; a fixed threshold misclassified over a fifth of successful projects. A size-scaled tolerance flagged a failing project before it failed.
D10 · Drift detection uses size-scaled thresholds -
DOI CHECKEDDEPENDENCIES
Hein, P. H., Kames, E., Chen, C., & Morkos, B. (2021). Employing machine learning techniques to assess requirement change volatility. Research in Engineering Design.
TL;DREach requirement is a multiplier, absorber, transmitter, or robust, and the class is predictable from network metrics on the requirements graph.
D6 · The dependency graph is a control surface -
DOI CHECKEDDEPENDENCIES
Arvanitou, E.-M., Ampatzoglou, A., Chatzigeorgiou, A., Avgeriou, P., & Tsiridis, N. (2022). A metric for quantifying the ripple effects among requirements. Software Quality Journal, 30(3), 853–883.
TL;DRA pairwise ripple probability between requirements, computed from past co-change and overlapping implementation.
D6 · The dependency graph is a control surface -
DOI CHECKEDSPEC SIZE
Ernst, N. A., & Robillard, M. P. (2023). A study of documentation for software architecture. Empirical Software Engineering, 28(5).
TL;DR65 newcomers answered architecture questions from one of two documentation formats. Format did not predict understanding; prior exposure to the source code did.
D2 · One source, two renderings -
DOI CHECKEDCONTROL THEORY
O'Neill, R. V., Johnson, A. R., & King, A. W. (1989). A hierarchical framework for the analysis of scale. Landscape Ecology, 3(3–4).
TL;DRA hierarchically organized system operates within a constraint envelope, with attractors inside it and thresholds where it changes state. Higher levels set boundary conditions on lower-level behavior.
D4 · Mandates carry numeric limitsD12 · Revocation is the kill switch -
DOI CHECKEDCONTROL THEORY
Holling, C. S. (1973). Resilience and stability of ecological systems. Annual Review of Ecology and Systematics, 4, 1–23.
TL;DRAn engineered device is judged by constancy around a goal. A system facing the unexpected is judged by persistence of relationships. The set-point tradition was carried into a domain where it may not apply.
D4 · Mandates carry numeric limits -
DOI CHECKEDOVERSIGHT
Zhu, L., Lu, Q., Ding, M., Lee, S. U., & Wang, C. (2026). Designing meaningful human oversight in AI. AI and Ethics, 6(3).
TL;DRShape the agent's output so a human can check it against an external criterion without redoing the work: structured rationales, policy attribution, circuit breakers, appeal bundles.
D3 · Rationale is a schema fieldD7 · Approve, deny, regenerate, edit; no silent edits -
DOI CHECKEDOVERSIGHT
Miller, T., Durlik, I., & Biczak, P. (2026). Adaptive human oversight for maritime agentic AI systems: Balancing operator workload and safety through risk-aware governance. Applied Sciences, 16(14), 6903.
TL;DRIn a maritime agentic-AI simulation, hysteresis and trend-aware cooldown cut repeated oversight requests and kept critical safety events visible.
D8 · Approvals are rationed and de-duplicated -
DOI CHECKEDOVERSIGHT
Chen, C., Zhang, Z., Chen, Z., Xu, E., Yang, Y., Khalilov, I., Gebreegziabher, S. A., Ye, Y., Xiao, Z., Yao, Y., Li, T., & Li, T. J.-J. (2026). Comparing human oversight strategies for computer-use agents. arXiv.
TL;DR4 oversight strategies, 48 participants, 192 live web sessions. Strategy changed how many problematic actions users met and left their ability to intervene unchanged. Plans constrained unplanned actions and left unlisted safeguards unaddressed.
D7 · Approve, deny, regenerate, edit; no silent edits -
DOI CHECKEDOVERSIGHT
Dhanorkar, S., Passi, S., & Vorvoreanu, M. (2026). Human oversight of agentic systems in practice: Examining the oversight work, challenges, and heuristics of developers using software agents. Proceedings of the 2026 ACM Conference on Fairness, Accountability, and Transparency (FAccT '26), 6438–6465.
TL;DRInterviews with 17 developers found 4 forms of oversight work: a priori control, co-planning, real-time monitoring, and post hoc review. Oversight is preventative and proactive as well as reactive.
D7 · Approve, deny, regenerate, edit; no silent edits -
DOI CHECKEDOVERSIGHT
Zhang, Y. E., & Wang, G. (2026). Towards human-centered agent authorization: A landscape analysis of commercial AI agents. Extended Abstracts of the 2026 CHI Conference on Human Factors in Computing Systems (CHI EA '26), 1–10.
TL;DR18 commercial agents surveyed. Bundled permissions, lazy disconnect, and non-persistent disclosure dominate. Revocation must be real.
D12 · Revocation is the kill switch
METHOD
Each entry gives the citation and DOI, the finding in one or two sentences, a quotation where the source was read, and the design decision it supports. A quotation means the sentence was read at the DOI or publisher page.
Compiled by Viberr · reviewed 14 Sep 2026 · corrections: ceo@viberr.me