Direct answer
The central enterprise AI question for 2026-2027 is no longer whether a model can generate plausible output. It is whether an organization can connect that output to reliable evidence, a controlled workflow, an accountable decision and a measurable result. Research should therefore study the whole operating system around AI, not model capability in isolation.
Evidence base
The NIST Generative AI Profile organizes risks and actions around governance, mapping, measurement and management. The European Union AI Act introduces obligations that vary by role and risk. Stanford's AI Index tracks rapid changes in capability, investment, use and governance. These sources are different kinds of evidence; none by itself proves that a particular deployment creates business value.
| Evidence signal | Management implication | Researchable question |
|---|---|---|
| General-purpose models are used across many workflows | Controls cannot be designed once for a single application | Which control patterns transfer across use cases without becoming empty checklists? |
| Risk obligations depend on context and role | Procurement needs technical, legal and operating evidence | What minimum evidence should be required before deployment? |
| Benchmark performance changes quickly | A static vendor score is insufficient | How often should a business-critical system be re-evaluated? |
| Value claims often rely on self-reported time saved | Productivity needs a baseline and quality measure | When does faster output improve throughput rather than create review work? |
Research streams
1. Decision-centered evaluation
Measure the decision outcome, not only model output. A useful study records the previous process, error cost, review time, escalation rate and downstream consequence. Comparison should include the status quo and a simpler non-AI intervention.
2. Human oversight as operating design
“Human in the loop” is not a control unless the person has time, authority, competence and evidence. Research should compare review designs: sampling, dual approval, threshold escalation, independent challenge and post-decision audit.
3. Retrieval, provenance and data boundaries
Generative systems often combine model knowledge with organizational sources. Studies should test whether users can identify the source, date and authority of material claims, and whether access controls survive indexing, retrieval and logging.
4. Agentic work
Systems that use tools and act across files or applications create a different control problem from chat. Research should examine permission boundaries, reversible actions, test environments, activity logs and the conditions that require explicit approval.
Proposed study design
Select one repeatable workflow with at least fifty historical cases. Define quality, time, cost and risk measures before introducing AI. Compare the current process, an AI-supported process and, where feasible, a process improvement without AI. Record exceptions and review work. Report confidence intervals or the full distribution instead of only an average.
Limits
This agenda does not estimate a universal AI return on investment. Results will vary by task, sector, data quality, worker experience, model version and control design. Legal obligations must be assessed for the actual jurisdiction and use.
Open materials
The CSV horizon scan records the evidence signals behind this agenda. The PDF methodology explains selection and limitations.
Editorially reviewed for clarity and source currency on by Igor Dmitriev, MBA, MsEM, MsIE .
