A Research Brief on the Economic, Business, and Accounting Implications of Effective Human-AI Conversation (2022-2026)
The launch of ChatGPT in November 2022 gave ordinary workers, students, and households a daily conversation partner for the first time. Five years on, that conversation has become a measurable economic activity. ChatGPT reached 100 million weekly active users within a year of launch and about 800 million by October 2025, while the share of US adults who say they have ever used it rose from 14 percent in March 2023 to 44 percent by February 2026 (Figures 1 and 2). Organisations moved even faster than consumers: the share of firms regularly using generative AI in at least one business function rose from 33 percent in McKinsey's 2023 survey wave to 65 percent in early 2024 and 79 percent in the 2025 wave (Figure 7).
This report asks what the skill behind those conversations consists of, how it should be measured, and what it is worth. The skill, which this report calls prompt literacy, has three parts: crafting a clear instruction, fixing a poor response, and knowing when the conversation is worth having. The evidence shows that prompt design materially changes model output: chain-of-thought prompting raised GSM8K accuracy from 17.9 to 58.1 percent in Wei et al. (2022), and zero-shot chain-of-thought raised MultiArith accuracy from 17.7 to 78.7 percent in Kojima et al. (2022) (Figure 3). Controlled experiments show large short-run productivity gains, from 40 percent faster writing with ChatGPT (Noy and Zhang 2023) to 15 percent higher customer support productivity (Brynjolfsson, Li, and Raymond 2025), but also a sharp risk: consultants working outside the AI capability frontier were 19 percentage points less likely to produce correct solutions (Dell'Acqua et al. 2026) (Figures 5 and 6).
The central question of the report is whether accounting, long called the language of business, can become the structured vocabulary of the AI era. The answer is treated as a hypothesis, not a conclusion. Three parts of the thesis have real evidence: structured reporting taxonomies are being rebuilt for AI (XBRL International's 2025 modernisation is explicitly designed to accelerate AI integration), assurance of AI outputs is being claimed as an accounting-led function (IAASB's September 2024 Technology Position, IFAC and IESBA guidance, and Big Four AI assurance offerings), and the UK's XBRL filing programmes have produced more than six million structured documents. Two parts remain untested: whether accounting vocabulary improves prompting in controlled experiments, and whether accountants outperform other professionals at prompt design or verification. No rigorous study of either question was found, and this report says so plainly rather than asserting an advantage.
For business and economics, the practical conclusion is that prompt literacy is a real but partly measurable skill. It shows up in job postings (prompt engineering mentions rose from 1,400 to nearly 6,300 between 2023 and 2024), in education (more than 327,000 learners in Vanderbilt's prompt engineering course alone), and in the way work is organised (75 percent of knowledge workers reported using generative AI at work in 2024). There is still no accepted test of the skill, surveys measure different things, and posting counts can reflect hype as much as genuine demand. The report separates sourced evidence from interpretation, and labels every number with its survey wave and source so that readers can check the claims themselves. All data, charts, and the replication script are open access.
Prompt literacy — The skill of crafting clear AI instructions, fixing poor responses, and judging when a conversation with an AI is worth having.
Chain-of-thought prompting — A prompting technique that asks the model to reason step by step, sharply improving performance on arithmetic and logic benchmarks.
Capability frontier — The boundary beyond which an AI system's raw capabilities cannot reliably complete a task; prompt skill matters most inside the frontier.
Assurance — The accounting and auditing discipline of providing verifiable, evidence-based confidence, applied in the report to AI-era information.
Accounting as language — The thesis that accounting, long the language of business, could become the standardised vocabulary for instructing and interrogating AI systems.
Before November 2022, conversing with a general-purpose language model was a specialist activity. After November 2022, it became an everyday one. ChatGPT was the first widely available system that let an ordinary person hold a natural-language conversation with a model that could write, summarise, reason, and code. The quality of that conversation depends on a human skill that had no meaningful pre-AI counterpart: the ability to say what one wants clearly enough for a machine to act on it, to recognise when the machine's answer is wrong, and to repair the request until the answer is right.
This report examines that skill, which it calls prompt literacy, over the period 2022 to 2026 in a global context, with pre-November-2022 material used only as baseline context. It asks three questions. First, what does the skill consist of, and how should it be defined and measured? Second, what is its economic and business value: does it show up in job markets, wages, productivity, and the way firms organise work? Third, and most distinctively, can accounting, long called the language of business, capture the trends of the AI era and become the new vocabulary of the AI era: a standardised, verifiable way of instructing and interrogating AI systems?
The skill itself has three components, which this report keeps separate throughout because they are different abilities and are measured by different evidence. The first is decomposing an intention into a structured, to-the-point instruction: breaking a vague goal into steps, constraints, and success criteria. The second is choosing accurate vocabulary and phrases so the model can interpret the instruction: naming entities, specifying formats, and giving context. The third is rapidly diagnosing, exploring, and fixing what goes wrong through iterative testing: reading a bad output, guessing why it failed, changing the prompt, and trying again.
Two mistakes are common in public discussion. One is to treat every improvement in AI-assisted work as evidence that prompt skill matters, when part of the gain comes from the model itself and from general AI adoption. The other is to treat prompt literacy as a fixed profession called prompt engineering, when the evidence suggests it is becoming a general workplace capability. The report's central task is to separate evidence from metaphor, and to test the accounting-as-language thesis rather than assert it. Sections 2 to 4 define the skill and the data; Sections 5 to 7 report what is measured; Section 8 tests the accounting thesis; Sections 9 and 10 set out scenarios and research gaps.
ChatGPT reached 100 million weekly active users by November 2023, 200 million by August 2024, was reported on track to reach 700 million by August 2025, and passed 800 million by October 2025 (Figure 1). These are platform announcements, not independent measurements, but they establish the scale of the phenomenon: hundreds of millions of people now write instructions to a language model every week.
Weekly active user figures as announced by OpenAI, November 2023 to October 2025. The August 2025 figure was announced as "on track" rather than achieved. Vertical dashed line: ChatGPT launch, November 2022. Source: author compilation from OpenAI announcements and press reporting (Business Insider, Axios, TechRepublic).
Adoption in the United States rose quickly but remains partial. The share of US adults who had ever used ChatGPT rose from 14 percent in March 2023 to 18 percent in July 2023, 23 percent in February 2024, 34 percent in the February-March 2025 wave, and 44 percent by February 2026 (Figure 2). In the 2025 wave, 58 percent of adults aged 18 to 29 had used it, well above the 34 percent all-adult share. The important caveat for this report is that ever-use is not prompt literacy. Many users ask simple questions; the survey evidence cannot tell us how many users decompose, debug, and verify. General AI adoption must therefore be kept separate from prompt-skill-specific effects throughout.
Share of US adults saying they have ever used ChatGPT, by Pew Research Center survey wave: March 2023 (14 percent), July 2023 (18 percent), February 2024 (23 percent), February-March 2025 (34 percent), February 2026 (44 percent). All observations are post-launch; the pre-November-2022 baseline is zero because the product did not exist. Source: Pew Research Center short reads and reports, 2023 to 2026.
The job market showed the skill component almost immediately. In the US job-posting data compiled by Lightcast for the Stanford AI Index, postings mentioning prompt engineering rose from 1,400 in 2023 to nearly 6,300 in 2024, while postings mentioning generative AI rose from 16,000 to more than 66,000 (Figure 4, Section 5). Education responded even faster: Vanderbilt's "Prompt Engineering for ChatGPT" course passed 327,000 learners, Google launched its five-step Prompting Essentials course in October 2024, and UNESCO published AI competency frameworks for students and teachers in September 2024 (Figure 11, Section 9). What had been a niche technique discussed in machine-learning papers became, within two years, a mainstream workplace and classroom topic.
This report treats the skill as a general capability rather than a specialist technique, and that choice matters for measurement. A specialist technique can be studied in isolation. A general literacy has to be defined, taught, and assessed against the background of everything else people do with language and work. The next section defines it.
Prompt literacy is defined here as the ability to achieve a desired outcome from a language model through written instruction, where the outcome is produced by the model but the instruction is designed, evaluated, and repaired by the human. The definition deliberately excludes model skill: it is a property of the person, not the system. It also excludes general AI literacy, which earlier literature defined before ChatGPT as knowledge about what AI is, what it can do, and how to use it responsibly (Long and Magerko 2020; Ng et al. 2021). Prompt literacy is narrower and more specific: it is about the quality of the instruction itself.
The construct has three components, measured by different indicators and never conflated in this report:
Prompt crafting. The ability to convert an intention into a structured instruction with the right vocabulary, context, constraints, and format. The strongest evidence that crafting matters comes from benchmarking studies that hold the model and task fixed and vary only the prompt. Chain-of-thought prompting, which asks the model to reason step by step, raised PaLM 540B accuracy on GSM8K from 17.9 to 58.1 percent (Wei et al. 2022). Zero-shot chain-of-thought raised MultiArith accuracy from 17.7 to 78.7 percent and GSM8K accuracy from 10.4 to 40.7 percent on the same model (Kojima et al. 2022) (Figure 3).
Accuracy on standard-prompting versus chain-of-thought prompting for the same model and task, as reported in Wei et al. (2022) and Kojima et al. (2022). The paired bars show that the prompt, not the model, drives the improvement. Source: arXiv:2201.11903 and arXiv:2205.11916.
Prompt debugging. The ability to diagnose why an output is wrong and repair the instruction. Direct measurement is scarce. The nearest large-scale evidence is negative: small changes in wording can break model behaviour. PromptRobust, a benchmark of 4,788 adversarial prompts spanning 8 tasks and 13 datasets, found that contemporary large language models are not robust to adversarial prompt variations (Zhu et al. 2023). Field research with non-expert users found that people tend to write task-specific, one-off prompts and struggle to design robust, generalisable strategies (Zamfirescu-Pereira et al. 2023). In other words, the debugging component exists and is hard, but no standardised test of it has been published.
Perceived value. What people think the skill is worth, measured by self-reported confidence, training, and use. This is the weakest type of evidence and is reported separately throughout. Survey data in this report (Pew, Microsoft-LinkedIn, McKinsey) measure adoption and attitudes, not skill; the report never treats perceived value as evidence of ability.
The measurement situation is clear: the field has benchmarks that hold the human fixed and vary the model (HELM, Liang et al. 2023, is the canonical example), and benchmarks that hold the model fixed and vary the prompt, but no validated instrument that measures a person's prompt literacy as a stable trait. Performance on standardised prompt-design tasks is the right idea, but no agreed task battery exists, and results across studies are weakly comparable. The Data and Method section reflects this: every dimension in Table 1 names an indicator, its primary source, and one main limitation.
The report covers seven dimensions. For each dimension, Table 1 names a single indicator, the primary source used, and one main limitation. The three components of prompt literacy (crafting, debugging, and perceived value) appear as separate rows where relevant, so they are never conflated.
| Dimension | Named indicator | Primary source | One main limitation |
|---|---|---|---|
| Emergence of human-AI conversation as a skill | Growth in prompting-related job postings; adoption rates of AI tools at work | Lightcast job postings for the Stanford AI Index; Pew, McKinsey, Microsoft-LinkedIn surveys | Posting counts can reflect hype rather than genuine skill demand |
| Defining and measuring the skill: prompt crafting | Performance on structured-instruction benchmarks with the model held fixed | Wei et al. 2022; Kojima et al. 2022 | Benchmarks measure prompt effects, not individual skill levels |
| Defining and measuring the skill: prompt debugging | Robustness under adversarial prompt variation; user strategy quality | Zhu et al. 2023 (PromptRobust); Zamfirescu-Pereira et al. 2023 | No validated test of debugging skill exists; proxies are ad hoc |
| Defining and measuring the skill: perceived value | Self-reported use, training, and confidence | Pew, McKinsey, Microsoft-LinkedIn surveys | Self-reports overstate ability and are not comparable across surveys |
| Labour market and economic value | Posting counts for prompt-related roles; productivity effects from experiments | Lightcast; Noy and Zhang 2023; Brynjolfsson, Li, and Raymond 2025; Dell'Acqua et al. 2026; Peng et al. 2023 | Short horizons, small samples, and selection effects; gains may not persist |
| Business and organisational implications | Enterprise AI adoption and training rates; share of firms with AI usage policies | McKinsey State of AI; World Economic Forum Future of Jobs 2025; Microsoft-LinkedIn Work Trend Index 2024 | Self-reported and inconsistent definitions across surveys |
| Accounting as the language of business in the AI era | Use of structured accounting vocabularies in AI interaction; volume of AI assurance activity and guidance | XBRL International; IFRS and IASB; IAASB; AICPA-CIMA; Big Four publications | Evidence is nascent; much of the thesis is conceptual and must be labelled argument |
| Verification, risk, and the audit mindset | Prevalence of verification practices; error rates in AI-assisted work; organisational AI controls | Dell'Acqua et al. 2026; IAASB and AICPA-CIMA guidance; enterprise governance reports | Verification behaviour is hard to observe directly; surveys may overstate good practice |
| Future trends and scenarios | Growth in curricula, credentials, and professional-body guidance; model interface capability | Vanderbilt; Google; UNESCO; World Economic Forum; model capability benchmarks | Inherently speculative; presented as scenarios with stated assumptions, not forecasts |
Two data conventions run through the report. First, survey waves are labelled exactly. McKinsey's figures are reported by wave (2023, early 2024, 2025), Pew's by field period (for example, February-March 2025 rather than "June 2025", which is when the short read was published), and Microsoft-LinkedIn's by the 2024 annual survey. Second, primary figures are preferred over secondary restatements. Where a secondary source restates a primary statistic, the variant is noted; the Stanford AI Index's reading of McKinsey survey data (55 percent of organisations using AI in 2023, 78 percent in 2024) is reported separately from McKinsey's own gen AI adoption series (33, 65, and 79 percent), because the two measure different things and are often conflated.
All underlying data are open access. The compiled CSV files listed in Section 11 (Data Availability and Replication) back every statistic quoted in this report, with commented headers recording each wave and source. README_methodology.txt describes compilation methods and limitations, and scripts/replicate.py reproduces every chart and prints every statistic from the CSV files. One genuinely open dataset, the Anthropic Economic Index interaction-type file on Hugging Face, is downloaded by the script with local caching as a verification check. Every number in this report is printed by the replication script; the report contains no number the script does not produce.
The economic value question has two parts: does prompt skill show up in hiring, and does it raise measured productivity? The hiring evidence is strong on volume and weak on wages. Lightcast's US job-posting data show that mentions of prompt engineering rose from 1,400 postings in 2023 to nearly 6,300 in 2024; mentions of generative AI as a skill rose from 16,000 to more than 66,000; and mentions of large language modeling rose from 5,000 to 20,000 (Figure 4). The counts are postings, not de-duplicated job titles, and they are demand signals that can reflect hype as much as verified skill requirements. No peer-reviewed estimate of a wage premium specifically for prompt literacy was found in the period covered; the report therefore does not claim one.
Job postings mentioning each skill, 2023 versus 2024. Logarithmic scale because posting volumes differ by more than an order of magnitude across skills. Postings are demand signals, not de-duplicated job titles. Source: Lightcast analysis for the Stanford AI Index 2025.
The productivity evidence is stronger and comes from four controlled experiments with different tasks, populations, and measurement conventions (Figure 5). Noy and Zhang (2023) randomly assigned 453 college-educated writing professionals to use ChatGPT for a writing task and found that the treatment group completed tasks 40 percent faster with quality rated 18 percent higher by independent evaluators. Brynjolfsson, Li, and Raymond (2025) studied 5,172 customer support agents and found a 15 percent average increase in productivity, measured as issues resolved per hour, with the largest gains for novice and low-experience workers. Dell'Acqua et al. (2026) ran a preregistered experiment with 758 management consultants at Boston Consulting Group and found that consultants inside the AI capability frontier completed 12.2 percent more tasks on average and completed them 25.1 percent more quickly, with significantly higher quality. Peng et al. (2023) randomly assigned 95 professional developers to use GitHub Copilot and found that the treatment group completed the task 55.8 percent faster.
Measured gains from four controlled experiments: time to complete (Noy and Zhang 2023), productivity as issues resolved per hour (Brynjolfsson, Li, and Raymond 2025), tasks completed (Dell'Acqua et al. 2026), and task speed (Peng et al. 2023). Each experiment used different tasks and outcome measures; the bars are not directly comparable magnitudes.
The same experiments measure the boundary of the skill. In Dell'Acqua et al. (2026), consultants asked to solve problems outside the AI capability frontier were 19 percentage points less likely to produce correct solutions when they had AI access than when they did not (Figure 6). The finding has two lessons. It shows that prompt literacy includes knowing when not to delegate, and it warns against extrapolating headline gains to all tasks. The experiments also share three limitations that the report states rather than hides: they run over days or weeks, not years; they use tasks chosen by the researchers; and the participants knew they were being studied. Whether the gains persist, generalise, or compound is not yet established by any published study in the period.
Change in the likelihood of producing a correct solution for tasks outside the AI capability frontier, comparing consultants with AI access to those without. Source: Dell'Acqua et al. (2026), Organization Science 37(2):403-423.
Two cautions apply to all of these numbers. First, these experiments measure AI assistance, not prompt skill specifically; part of the gain would accrue to a user with a mediocre prompt. Second, the posting evidence measures demand for skills as advertised, which can lag or exaggerate the skills actually used on the job. The interpretation supported by the evidence is that prompt literacy raises the return to AI access for tasks inside the capability frontier, especially for less experienced workers, and that the return turns negative outside it. The evidence does not yet support a claim about the long-run wage premium for the skill, and this report does not make one.
Firms moved from experimentation to routine use within three years. In McKinsey's Global Survey on AI, the share of organisations regularly using generative AI in at least one business function rose from 33 percent in the 2023 wave to 65 percent in the early 2024 wave (fielded February 22 to March 5, 2024) and 79 percent in the 2025 wave (fielded June 25 to July 29, 2025); in the 2025 wave, 88 percent of organisations reported using AI in at least one function (Figure 7). The Stanford AI Index's reading of the same McKinsey data gives 55 percent of organisations reporting AI use in 2023 and 78 percent in 2024; the two series are kept separate because they measure different things.
Share of organisations regularly using generative AI in at least one business function, by McKinsey survey wave: 33 percent (2023), 65 percent (early 2024), 79 percent (2025). The final bar shows the 2025 wave reading of AI use in any function (88 percent). Source: McKinsey State of AI reports, 2023 to 2025.
At the worker level, Microsoft and LinkedIn's 2024 Work Trend Index, a survey of 31,000 people in 31 countries, found that 75 percent of knowledge workers use generative AI at work, that 66 percent of leaders would not hire someone without AI skills, and that 71 percent of leaders would rather hire a less experienced candidate with AI skills than a more experienced candidate without them. Pew's narrower measure of actual use paints a more sober picture: 16 percent of US workers said at least some of their work was done with AI in the October 2024 survey, rising to 21 percent in the September 2025 survey (Figure 8). The gap between the two surveys is partly definitional (knowledge workers versus all workers; "use" versus "at least some of my work is done with AI") and partly real; this report shows both rather than averaging them.
Left pair: share of US workers saying at least some of their work is done with AI, October 2024 (16 percent) and September 2025 (21 percent). Right group: Microsoft and LinkedIn Work Trend Index 2024 readings for knowledge workers (75 percent use generative AI), leaders who would not hire without AI skills (66 percent), and leaders preferring a less experienced candidate with AI skills (71 percent). Definitions differ across surveys and are not directly comparable.
Platform evidence shows how AI is actually woven into work. Anthropic's Economic Index, based on anonymised Claude conversations, found that about 36 percent of occupations used AI for at least a quarter of their associated tasks in the January 2025 sample, with about 4 percent using it across three-quarters or more of tasks; pooling data across reports, the share of occupations reached 49 percent in the January 2026 report (Figure 9a). The first report classified 57 percent of AI tasks as augmentation, where AI collaborates with the worker, and 43 percent as automation, where AI directly performs the task; by the November 2025 sample, augmentation stood at 52 percent and automation at 45 percent (Figure 9b). The pattern that holds across both reports is that AI use is broad but shallow: many occupations touch AI on some tasks, very few delegate most of their work to it.
Share of occupations with Claude used in at least 25 percent of associated tasks: January 2025 sample (36 percent) and November 2025 sample pooled with earlier reports (49 percent). Source: Anthropic Economic Index first report (10 February 2025) and fourth report (15 January 2026).
Share of Claude conversations classified as augmentation or automation: January 2025 sample (57 percent augmentation, 43 percent automation) and November 2025 sample (52 percent augmentation, 45 percent automation). Source: Anthropic Economic Index first and fourth reports.
Three organisational implications follow. First, the roles changed faster than the skills: prompt engineering and AI trainer roles appeared in job postings (Section 5), and Microsoft-LinkedIn's hiring numbers show that employers expect AI skills at entry. Second, training became a board-level question: the World Economic Forum's Future of Jobs Report 2025 expects 39 percent of workers' core skills to change by 2030, and 85 percent of employers say they plan to prioritise reskilling and upskilling. Third, governance lagged adoption: the survey evidence measures usage far better than it measures usage policies, and the report flags that gap rather than filling it with anecdotes. The self-reported nature of all these surveys is their shared limitation; definitions of "regular use", "use at work", and "AI skills" differ across McKinsey, Pew, and Microsoft-LinkedIn, and cross-survey comparisons are indicative only.
Effective AI conversation does not end when the model replies. The third component of the skill is knowing how to check outputs, catch errors, and recognise when not to trust the model. The evidence that this matters comes from the same experiment: the study that showed 12.2 percent more tasks completed inside the frontier also showed a 19 percentage point drop in correct solutions outside it (Dell'Acqua et al. 2026, Figure 6). If users cannot tell which side of the frontier they are on, the expected value of AI assistance can be negative.
Direct evidence on verification behaviour is thin, and the report says so. Verification is hard to observe: a user who quietly fact-checks an output leaves no trace in most datasets, and surveys asking people whether they verify outputs may get flattering answers. PromptRobust's finding that 4,788 adversarial prompts across 8 tasks and 13 datasets degrade model performance (Zhu et al. 2023) implies that users face an unstable target: the same request can succeed or fail depending on small wording changes. Anthropic's own data show that users treat AI outputs differently by task, with augmentation (where the human checks and edits) remaining the most common pattern at 52 percent of conversations in the November 2025 sample, against 45 percent automation.
The audit mindset hypothesis is that accounting-trained people, whose professional discipline is built on verification, documentation, and professional scepticism, are better positioned for this component of the skill. That hypothesis is plausible and currently unsupported by direct evidence. No published study comparing accountants with other professionals on prompt design or verification tasks was found in the period covered, and the report does not claim an advantage. What the evidence does support is narrower: organisations are building verification into AI governance (IAASB's Technology Position and the assurance guidance discussed in Section 8), and the market for checking AI outputs is emerging as an accounting-adjacent service. Section 8 examines that development directly.
Accounting has long been called the language of business because it gives economic activity a standardised vocabulary: chart of accounts, double-entry logic, accruals, disclosures, and audit verification. The thesis examined here is that this structured vocabulary can become the new vocabulary of the AI era: a standardised, verifiable way of instructing and interrogating AI systems for business decisions. The thesis has three testable parts, and the report's discipline is to state, for each part, what is evidence, what is argument, and what is absent.
Part one: structured accounting vocabularies improve human-AI communication. The strongest evidence is infrastructure-level. XBRL International announced in March 2025 a modernised taxonomy specification explicitly designed to "directly facilitate and accelerate the integration of artificial intelligence with structured data and business rules", and followed with publications on making data AI-ready (October and November 2025). The UK's XBRL filing programmes, which require company tax filings and voluntary accounts filings in structured form, have produced more than six million XBRL documents. This is evidence that structured business reporting is being rebuilt for machine consumption, and that accounting-standard vocabulary is becoming machine-readable. It is not evidence that accountants write better prompts; the step from structured taxonomies to better instructions is an argument from analogy that no published experiment has tested.
Part two: accountants perform differently from others on prompt-design tasks. No rigorous study of this question was found. The report searched for controlled comparisons of accountants versus other professionals on prompt crafting, debugging, or verification, and found none in the period covered. There is academic work on how well language models perform on accounting exam tasks, but that measures the model, not the accountant, and is not cited as evidence about accountants' prompting skill. Part two is therefore recorded as untested: a clear research gap, not a finding.
Part three: assurance of AI outputs is emerging as an accounting-led function. This has the clearest institutional evidence. IAASB adopted its Technology Position at its September 2024 meeting, setting out how audit and assurance standards will adapt to technology. IFAC and IESBA published "Navigating the Gen AI Revolution: Implications for the Accounting Profession" in June 2024. In June 2025, Deloitte, EY, and PwC were reported as preparing AI assurance services, and PwC launched an AI assurance offering that month; these launch items are media-reported rather than primary firm publications and are labelled as such. Figure 10 assembles the timeline.
Timeline of standard-setting and firm activity: IFAC and IESBA guidance (June 2024), IAASB Technology Position (September 2024), XBRL International taxonomy modernisation (March 2025), Big Four AI assurance preparation as reported by the Financial Times and City AM (June 2025), PwC's AI assurance launch (June 2025), XBRL International publications on AI-ready data (October and November 2025), and the UK's six million XBRL documents (accessed 2026). Big Four launch items are media-reported.
The three parts of the thesis get three different answers. Part three is evidenced: the standard-setting bodies and the largest firms are moving, in their own publications, toward AI assurance as a professional function. Part one is partially evidenced at the infrastructure level and argued at the skill level. Part two is untested. Overall, the thesis that accounting can become the language of the AI era remains a hypothesis: institutions are moving toward AI assurance, but no direct evidence shows that accounting training improves the human side of the conversation. The report also notes what it could not verify: claims that accountants are better prompters, and specific claims about Big Four AI assurance market size, are either unsupported or media-reported, and are not repeated as fact.
Whether prompt literacy becomes a general literacy, a specialist profession, or a transitional skill is a question about the future, and this report answers it with scenarios, each with stated assumptions, rather than point forecasts. The leading indicators are the education response (Figure 11) and the pace of interface improvement.
Timeline of education responses: UNESCO AI competency frameworks for students and teachers (September 2024), Google Prompting Essentials launch (October 2024), Vanderbilt reporting more than 327,000 learners in its prompt engineering course and 500,000 across related courses (October 2024), and the World Economic Forum Future of Jobs 2025 expectations that 39 percent of core skills will change by 2030 and 85 percent of employers will prioritise reskilling (January 2025).
Scenario one: general literacy. Prompt literacy becomes a basic capability like reading and writing, embedded in school curricula (UNESCO's frameworks point this way) and expected of every knowledge worker (Microsoft-LinkedIn's 66 percent of leaders who would not hire without AI skills point this way). Under this scenario, the accounting implication is that structured vocabulary matters more, not less: as everyone learns to converse with AI, the professional premium shifts to domain vocabularies and verification discipline, which is exactly the accounting thesis. The economic implication is a broad, mostly positive productivity shift with distributional costs for workers who do not acquire the skill.
Scenario two: specialist profession. Prompt engineering becomes a certified specialism, as the first wave of job postings and courses suggested (1,400 to nearly 6,300 postings from 2023 to 2024; 327,000 learners by October 2024). Under this scenario, organisations hire prompt engineers and AI trainers, and professional bodies develop credentials. The accounting implication is that the AI assurance function becomes a distinct specialism alongside audit, with its own standards and liability questions. The economic implication is concentrated rents for a small group and slower diffusion of the skill.
Scenario three: transitional skill. Interfaces improve faster than skills diffuse, so the explicit craft of prompting fades into ordinary interaction: models become better at inferring intent, handling ambiguity, and checking their own work. Under this scenario, today's prompt engineering premium decays, and the durable skill is the underlying ability to specify problems and verify answers, which is the same decomposition and verification discipline the report identifies at the core of prompt literacy. The accounting implication is that the language-of-business thesis survives in a stronger form: structured specification and verification remain valuable even when the prompting craft disappears. The economic implication is that today's measured productivity gains understate the long-run effect, because the skill requirement falls while the capability rises.
The report does not weight the scenarios; their likelihood depends on how fast interfaces improve, which is the hardest input to predict. One conclusion holds across all three scenarios: the underlying abilities, decomposing problems, specifying them precisely, and verifying outputs, are the core of the construct.
The evidence reviewed points to eight research opportunities, rated in Figure 12 by theory significance and data availability on 0 to 10 scales. The ratings are the author's analytical assessment, not survey measurements, and are stated as such in the data file.
Author ratings (0 to 10) of eight research opportunities. Axes are zoomed to the range the data occupy: data availability 2 to 7, theory significance 6 to 9. The ratings are analytical judgements recorded in research_opportunities.csv, not survey or experimental measurements.
Four gaps matter most for economics, business, and accounting scholars. First, the wage premium question: with posting data showing demand and experiments showing short-run productivity, the missing piece is a wage estimate that links verified prompt skill (rather than job titles) to earnings. Second, persistence: every headline experiment is short-horizon, and no study in the period shows whether the 40 percent writing gain or the 15 percent support gain survives a year of routine use. Third, measurement: the field needs a validated prompt literacy assessment, without which the construct cannot be compared across studies, occupations, or countries. Fourth, the accounting question: controlled experiments comparing accountants with other professionals on prompt crafting, debugging, and verification would test the language-of-business thesis directly, and none exists yet.
Two more gaps are worth noting. Organisation-level training return on investment is the question boards actually face, and the data are proprietary and inconsistent. And the AI assurance market is being built in real time, but its size, pricing, and quality are not yet measured in any public dataset. Each of these is a realistic project for a dissertation or a grant, which is why they are rated explicitly.
All data, methodology, and code are open access and downloadable. The compiled datasets back every statistic in this report, and each CSV begins with commented headers naming the indicator, wave, and source.
The replication script reproduces every chart in this report as a PNG and prints every statistic quoted in the report and summary. Its printed output is the source of truth: the report contains no number the script does not produce. The script auto-downloads the open Anthropic Economic Index interaction-type dataset with local caching, and reads all other datasets from the data folder because the underlying organisations publish headline figures in PDFs and press releases without stable machine-readable endpoints. The script sets an explicit random seed (42), pins dependency versions, checks that every data file and column exists, and organises its output by report section.
Primary sources for the statistics quoted in this report. Every URL below was fetched or searched before publication; publisher blocks (403) from academic sites are noted where they occur.
Academic DOIs were verified against Crossref (title, journal, volume, pages, year) before publication. Publisher sites may return 403 to automated requests; none of the cited URLs returned 404 at verification time.