Characterizing Agents in Production
SAI paper + code review · Referee report
Summary
MAP (Measuring Agents in Production) reports the first systematic, cross-organization empirical characterization of LLM-based agents that have actually reached production, drawing on 20 in-depth practitioner interviews and a filtered subset of 86 deployed systems from a 306-respondent survey collected between April and November 2025. The conceptual move — trading depth of any one deployment for breadth across sectors, geographies and organizational maturities — is genuinely novel: prior work is either single-system engineering blog posts, commercial-consulting adoption surveys, or academic literature reviews, and none has attempted to elicit engineering-level design decisions across dozens of live systems. The paper's central substantive claim is that production teams systematically choose the "boring" option — prompted off-the-shelf frontier models, static workflows with ≤10 autonomous steps, custom in-house scaffolds, and human-in-the-loop evaluation — not from methodological naivety but as a deliberate reliability strategy that trades capability for controllability. That framing (production simplicity as a design principle, not a transitional state) is a useful counterweight to the research literature's focus on autonomy and RL post-training, and the 13 findings collectively form a plausible ontology of what "deployment realities" currently look like. The main conceptual limitation is that the study samples only deployments that survived to production, so every reported pattern is confounded with a survival filter; the paper cannot causally attribute successful deployment to any of the practices it observes, and it does not sample failed or abandoned agent projects. A second limitation is that all quantitative claims are self-reported and un-audited, with no raw data release, so cross-tabulations that would test the paper's own conjectures (latency × user-type, autonomy × domain, framework × maturity) cannot be run internally or externally. Within those caveats the paper is a valuable community resource and the qualitative themes are likely to hold up.
Strengths
- Genuinely novel data. MAP is the only work we are aware of that combines depth (20 semi-structured interviews with hands-on engineers) with breadth (86 deployed systems across 26 self-reported domains) and reports engineering-level decisions rather than executive perceptions. The abstract's emphasis on practitioner data is well-earned.
- Balanced framing of the research-practice gap. The discussion is careful to name both confirmatory findings (productivity is the primary motivation) and counterintuitive ones (bounded autonomy, near-total reliance on prompting, minute-scale latency tolerance), rather than only emphasizing the surprises. Section 8 usefully maps these back to concrete research opportunities (sample-efficient post-training, evolutionary search, observability, coordination interfaces).
- Methodological transparency about qualitative process. The paper documents the interview protocol, grounded-theory open coding, three-annotator consensus with κ=0.636 and adjudication, and the semi-structured protocol's 11 topic areas. This is unusually explicit relative to comparable consulting reports.
- Useful data-driven challenges to research narratives. The observations that (i) latency-sensitive deployments are the minority (Fig. 4), (ii) practitioners deliberately migrate from CrewAI-style frameworks to custom stacks, and (iii) SFT/RL is discouraged specifically because it is brittle to model provider upgrades, are concrete and actionable for the research community and are supported by the interview evidence.
Weaknesses
- Selection bias and no failure comparison group. The study samples deployments that survived to production/pilot, so the reported patterns cannot be causally attributed to deployment success — they describe what surviving systems look like. Given the paper's own framing of a ~40% agent-project failure rate (Reuters/Gartner), the absence of failed or abandoned projects means observations like "practitioners favor simplicity" could equally be read as "surviving projects happen to be the simple ones". This is the paper's most consequential methodological limitation and is not adequately foregrounded.
- Snowball recruitment via authors' networks. All recruitment channels (Berkeley RDI Summit, AI Alliance Meetup, Berkeley Agentic AI MOOC, LinkedIn/Discord/X) are proximate to UC Berkeley RDI and the authors' professional networks. Snowball case-study sampling further concentrates this bias. The Limitations paragraph acknowledges "participation bias" but does not name a plausible direction (over-representation of ML-mature, Bay-Area-linked organizations that already identify as building "agents"); the prevalence estimates should be read with this bias in mind.
- Headline percentages rest on per-question denominators much smaller than the "86 deployed" framing. The 68% autonomy figure is 41/60 (Figure 7a), the 74% human-in-the-loop figure is 23/31 (Figure 8), the 80% productivity figure is 53/66 (Figure 1), the 66% latency figure is 35/53 (Figure 4). Because "all questions were optional", item-level non-response is 22-64% of the 86 and is not itemized; the responding subset may not be representative of the deployed sample.
- Framework finding cherry-picks case-study evidence. Finding 10 asserts "Production agents favor custom implementations (85%) over frameworks" but the survey (61% use frameworks) points the opposite direction. The 85% custom finding may hold specifically for mature engineering organizations that dominate the case studies (Fig. 15: 15/20 late-stage or mature), but as written the headline over-generalizes from case studies to "production agents" as a class.
- Self-report and confidentiality without independent verification. All measurements — "10x faster than humans", "<10 autonomous steps", "read-only agents", counts of models per system — are practitioner-reported and confidential. There is no attempt to audit any single deployment against actual system behavior. Social-desirability bias is not mentioned. The paper cites direct auditing "as follow-up work" but writes the current findings with the confidence of measured facts.
- Reliability is defined formally but measured coarsely. Section 7 defines reliability using the IEEE probability-of-failure-free-operation frame, but the actual empirical evidence is a Rank-1 vote for a bundled category ("Core Technical Performance": robustness, reliability, scalability, latency, resource constraints) chosen by 38% (11/29). The claim that reliability is the "top development challenge" is therefore weaker than the definitional framing suggests.
- Long tail of 26 domains rests on singleton responses. Beyond the top three sectors, most reported domains have 1-3 responses each, and Table 1 lists 16 "Other" domains with n=1. The strong "26 domains suggest expanding opportunities" reading is not well supported by n=1 evidence.
- Conjectures are presented as correlations without cross-tabulation. Section B.4.2 juxtaposes latency tolerance (15/20 async) with internal user prevalence (52.2%) and infers they "correlate", but the two come from independent marginal survey items and are never cross-tabulated. Similar juxtapositions appear when linking security posture to internal-employee share and productivity focus to verification difficulty. The raw data would support these tests; the paper leaves them un-run.
- Several internal-consistency issues in numbers and figures. Notable: (a) Abstract says "surveyed 86 deployed systems practitioners" but 306 practitioners were surveyed and 86 deployed systems were filtered; (b) Figure 1 caption calls the intervals "95th percentile intervals" instead of 95% CIs; (c) Finding 5 says 3/20 open-source but Figure 5 shows 3 "Yes" + 3 "Sometimes" = 6/20; (d) C.1 states 82% production/pilot but 45.0% + 32.4% = 77.4%; (e) Figure 21 caption's rationale attaches the "stricter control" reasoning to the wrong subset; (f) Figure 25 caption contradicts the body about which subset shows the stronger non-text focus; (g) Table 3 "Software & Business Operations" contains C14-C17 but the prose calls it C14-C16; (h) the "aspect → survey question" mapping table lists QN7 as "Metrics" and QN3 under "Application domain" and "System stage", none of which match those questions.
- Survey instrument has ambiguities. QN5.1's user-count bins ("1-10" then "10-49") overlap at 10, and QN10's autonomy bins offer "Four or fewer" and "Ten or fewer" as adjacent MCSA options that both apply to a 3-step system. QN14 offers 11 latency options collapsed into 6 display bins without a documented mapping. These would be less concerning if the raw data were released, but combined with non-release they mean several bin counts are non-recoverable.
Reproducibility & code
The paper claims data-driven insights but does not accompany the manuscript with a data or code release. Reading the supplied materials (paper.md and its appendices) surfaces the following coverage gaps; nothing was executed.
- No raw survey response data. None of the 306 (or 86 deployed) per-respondent answers to QN1-QN14 / QO1-QO18 is released. Every headline percentage (68%, 70%, 74%, 80%, 66% ...) is therefore a "trust us" number: figure bin counts sum correctly, but the counts themselves cannot be checked against a raw response CSV.
- No case-level coding table for the 20 interviews. All statements of the form "17/20 use closed-source", "14/20 no post-training", "16/20 structured workflow", "15/20 async", "3/20 use frameworks" rest on interview coding that is never published. Table 3 further anonymises 20 cases as 17 IDs (C01-C17) by merging similar cases, without documenting the merge rule, so even the anonymised identifiers cannot be mapped 1-to-1 back to case-study statistics.
- LOTUS domain-normalization pipeline is under-documented. Appendix B.1.2 shows a single LOTUS prompt but does not pin the LOTUS version or the underlying LLM, and does not release the raw QN4 free-text answers, the per-annotator labels for the three annotators, or the adjudicator's final labels. Cohen's κ=0.636 cannot be recomputed externally, and the "26 domains" ontology cannot be re-derived.
- Bin-collapse rules are not documented. QN10 offers 7 autonomy options (including two overlapping ones) that Figure 7a maps to 6 bins; QN14 offers 11 latency options that Figure 4 maps to 6 bins. The mapping rules — especially how the overlapping QN10 options were resolved and how Weeks/Months/More were folded into ">1 day" or "No limit" — are not stated, so a reader with the raw data could not exactly reproduce the figure bins.
- Bootstrap CI procedure is described but not re-runnable. The paper says "95% CIs from 1,000 bootstrap samples with replacement", but no seed, implementation, or resampling notebook is released; combined with no raw data, the reported CIs are not verifiable.
Recommended Changes
Essential
- Reframe the study as descriptive of survived deployments. In the abstract, intro and discussion, make explicit that the study samples only deployments that reached production/pilot and therefore cannot causally distinguish which practices caused success. Add a limitations paragraph naming this survival filter and note the absence of failed/abandoned projects. (Weaknesses: selection bias / no failure comparison.)
- Fix the abstract's "surveyed 86" phrasing. Rewrite the first substantive sentence to something like "surveyed 306 practitioners, of whom 86 reported systems in production or pilot phases", so the survey scope is not misstated. (Weaknesses: internal-consistency issues; inline comment on abstract wording.)
- Report per-question denominators alongside every headline percentage. In the abstract and executive-summary paragraphs, state N alongside each of 68%, 70%, 74%, 80%, 66%; add item-level non-response to Section 3.2; and provide a sensitivity analysis or CI on the headline figures under plausible non-response scenarios. (Weaknesses: headline percentages rest on smaller denominators.)
- Release an anonymised response CSV, the case-level coding matrix, the LOTUS version/prompt, and the analysis notebook. These would resolve most of the reproducibility gaps above (raw data, LOTUS pipeline, case-level coding, bin-collapse rules, bootstrap CIs). At minimum, publish the case-level coding table with anonymised IDs kept 1-to-1 with the 20 cases (not merged into 17). (Reproducibility & code: all bullets.)
- Reconcile Finding 10's 85%-custom claim with the 61% survey framework use. Either explain the gap (e.g. as a maturity confound observable in Fig. 15) or hedge Finding 10 to say custom stacks dominate the mature case-study organisations specifically, not "production agents" in general. (Weaknesses: framework finding cherry-picks evidence.)
- Correct the numerical inconsistencies. Update Finding 5 to reflect 6/20 (Yes + Sometimes) open-source use per Figure 5; correct "82% in production or pilot" to ~77%; align Table 3's "C14-C17" group with the prose narrative; fix Figure 21's inverted rationale; fix Figure 25's caption vs body contradiction; align the Figure 1 caption's interval terminology with Section 3.2 ("95% CI"); fix the "Figure ??" cross-references. (Weaknesses: internal-consistency issues.)
Suggested
- Add an explicit self-report/audit caveat. In Section 8, note that all measurements are practitioner-reported perceptions and counts, and that direct auditing of deployed systems is deferred as follow-up. Downgrade point-precision claims like "10x faster than humans" and "68% execute <10 steps" to the qualitative claim they can support. (Weaknesses: self-report bias; inline comment on 10x.)
- Run and report cross-tabulations you already have data for. Latency-tolerance × user-type, human-in-the-loop × domain, framework-use × organizational maturity, autonomy step count × sector — each would test conjectures the paper currently juxtaposes as "correlations". Move Section B.4.2's latency conjecture and the security-by-internal-employees conjecture behind those cross-tabs. (Weaknesses: conjectures presented as correlations.)
- Separate reliability from the bundled "Core Technical Performance" category. Either report a reliability-specific question if one exists in the instrument, or reframe Section 7.1 to say "the bundled category dominated by reliability" rather than "reliability remains the primary bottleneck". (Weaknesses: reliability measurement.)
- Add a mapping table from the 13 findings to the 3 discussion themes. Explicitly assign F1-F13 to the three themes in Section 8 (or explain unassigned findings), so the "13 → 3" consolidation is verifiable. (Weaknesses: consolidation traceability; inline comment on F→theme mapping.)
- Audit the aspect-to-question mapping table. Correct the mislabelling of QN7 as "Metrics", the reuse of QN3 under both "Application domain" and "System stage", and duplicate QN9/QN10/QN11 references across control-flow and dependencies aspects. (Weaknesses: internal-consistency issues; inline comment on aspect mapping.)
- Restrict "26 domains" language to sectors with adequate sample size. Rephrase Finding 2 and the diversity claim so that n=1 domains from Table 1 are described as "long-tail one-off responses", not evidence of a prevalence pattern. (Weaknesses: long-tail domain claims.)
- Disambiguate survey instrument boundaries. Fix overlapping bins in QN5.1 (1-10 / 10-49) and QN10 (Four or fewer / Ten or fewer), document the QN14→Figure 4 bin-collapse rules, and clarify whether "end users" in Figure 14c are daily or cumulative. (Weaknesses: instrument ambiguities.)