The Curated Future Brief: What the Numbers Say About AI Today
AI is becoming cheaper to use, more expensive to build, and harder to measure by any single benchmark. A field guide to the figures shaping products, culture, labor, and the next generation of startups.
Marek DvořákSenior product reviewerFirst published 8/21/2026 · monitored for updates; the next revision publishes a new version and appears here. Reader corrections are reviewed and folded into future versions.
Summary
AI’s defining contradiction is visible in the numbers: frontier models require extraordinary capital and infrastructure, while intelligence delivered through an API is rapidly becoming cheaper and more accessible. Stanford’s 2025 AI Index estimated that inference costs for performance comparable to GPT-3.5 fell more than 280-fold between November 2022 and October 2024, even as industry produced nearly 90% of notable models in 2024. Adoption is broad, but value is uneven: many organizations use AI, fewer have redesigned workflows around it, and even fewer can quantify durable returns. For founders and creative strategists, the opportunity is shifting away from generic generation toward trusted interfaces, proprietary context, elegant workflow design, and products that make imperfect intelligence genuinely useful.
Key takeaways
- Stanford reported that 78% of surveyed organizations used AI in 2024, up from 55% in 2023; experimentation has become mainstream.
- The cost of querying capable models is collapsing, inviting AI into low-margin and high-volume products that could not previously support it.
- Frontier development is concentrating: industry produced nearly 90% of notable AI models in 2024, reflecting the capital intensity of compute, data, and talent.
- Performance gaps between leading models are narrowing, so distribution, product taste, latency, trust, and proprietary data increasingly determine advantage.
- Generative AI attracted $33.9 billion in global private investment during 2024, according to Stanford—strong conviction, but not proof of widespread profitability.
- The labor story is more often task redesign than instant job replacement: exposure varies sharply by occupation, language, firm, and workflow.
- AI’s physical footprint matters. Data-center electricity demand is rising, while local inference and more efficient models offer partial counterweights.
- Benchmarks describe controlled performance, not complete products; real-world quality also depends on retrieval, evaluation, safeguards, interaction design, and human judgment.
Explain like I'm 5
Imagine AI as a very talented, extremely fast apprentice. It has read an enormous library, can sketch, code, summarize, and imitate many styles—but it can also confidently invent details. The newest apprentices are costly to train because they need warehouses of specialized chips and electricity. Once trained, however, hiring one for a small digital task is becoming dramatically cheaper. That creates two different markets. A small number of companies build the underlying intelligence; thousands of others package it into useful experiences. The winning package is rarely just a chat box. It remembers the right context, fits naturally into a task, checks its own work, shows uncertainty, protects private information, and knows when to ask a person. The important number, therefore, is not simply a benchmark score. It is the measured improvement in a real outcome: hours saved, errors reduced, concepts explored, customers retained, or work made possible.
Deep dive
Adoption is broad; transformation is scarce
Stanford’s 2025 AI Index reported that 78% of surveyed organizations used AI in at least one business function in 2024, compared with 55% one year earlier. McKinsey’s 2025 global survey similarly found widespread regular use, with generative AI especially common in marketing and sales, product and service development, software engineering, and IT. Yet adoption is a soft metric: opening a Copilot seat and rebuilding a company around machine-assisted work both count as use. The more revealing questions concern frequency, workflow penetration, error rates, unit economics, and whether employees keep returning after novelty fades. For builders, this gap is fertile ground. Products that capture context, orchestrate approvals, preserve provenance, and fit existing tools can turn casual experimentation into dependable infrastructure.
Intelligence is becoming a cheaper material
Model capability has improved while the price of accessing it has fallen unusually fast. Stanford estimated that the inference cost of achieving GPT-3.5-level performance on the MMLU benchmark dropped more than 280-fold from November 2022 to October 2024. Small models also became substantially stronger, and open-weight systems narrowed portions of the gap with closed leaders. This changes product arithmetic: classification, transcription, translation, image understanding, and personalized generation can enter services with modest revenue per user. But cheap tokens do not guarantee a cheap system. Retrieval pipelines, evaluation, observability, moderation, human review, and support often dominate production costs. The design brief is no longer ‘add AI’; it is to spend intelligence selectively, routing easy requests to small models and difficult ones to more capable systems.
The frontier remains expensive and concentrated
Falling inference prices coexist with rising development costs. Stanford’s estimates place the training compute cost of several recent frontier models in the tens or hundreds of millions of dollars, before fully accounting for research salaries, failed runs, data work, and deployment infrastructure. Industry created nearly 90% of notable AI models in 2024, while academia remained important in highly cited research. This concentration favors laboratories with chip access, cloud capacity, distribution, and cash. It also clarifies where most startups should not compete. Training a general-purpose frontier model is usually a weaker proposition than owning scarce domain data, a trusted audience, a regulated workflow, or a distinctive interface built across several model providers.
Work changes at the level of tasks
The International Monetary Fund estimated in 2024 that almost 40% of global employment is exposed to AI, rising to about 60% in advanced economies. Exposure is not equivalent to elimination. Studies of customer support, consulting, and software development have found productivity gains in bounded settings, often with larger benefits for less-experienced workers; other research finds mistakes, uneven uptake, or slower performance when tools are poorly matched to a task. Generative systems are particularly good at producing plausible first drafts and variations. Humans remain essential where the work requires responsibility, physical action, tacit context, negotiation, original direction, or a defensible claim that something is true. Creative roles may become less about making the first artifact and more about framing, selecting, editing, and establishing authorship.
The cultural supply shock
Text, imagery, music, and video are becoming abundant at a speed that unsettles existing ideas of craft and value. Adobe, Canva, Runway, Midjourney, OpenAI, Google, and others have compressed sophisticated production into prompts and lightweight interfaces. The consequence is not merely more content; it is a change in what audiences reward. When competent output is plentiful, discernment, specificity, live experience, provenance, and human narrative can become more valuable. Copyright disputes—including cases involving The New York Times, Getty Images, artists, record labels, and model developers—show that training data is an economic and cultural question, not a technical footnote. Products that attach permissions, attribution, compensation, and editable provenance to generated media may become important creative infrastructure.
Energy is part of the interface
The International Energy Agency estimated that data centers consumed roughly 415 terawatt-hours of electricity in 2024, about 1.5% of global electricity use, and projected demand could more than double to around 945 TWh by 2030 in its base case. AI is not the only workload in that total, but it is a major source of growth. Aggregate figures also hide local pressure: a new campus can strain a particular grid, water system, or permitting regime. Efficiency gains are real—better chips, quantization, distillation, and local models reduce energy per task—but lower prices can stimulate much greater use. Responsible product strategy should therefore measure quality per unit of compute, not celebrate model size as an end in itself.
- 1956The Dartmouth workshop, organized by John McCarthy, Marvin Minsky, Nathaniel Rochester, and Claude Shannon, popularizes the term ‘artificial intelligence.’
- 1997IBM Deep Blue defeats world chess champion Garry Kasparov in a six-game match, making machine specialization culturally visible.
- 2012AlexNet wins the ImageNet competition by a wide margin, accelerating industry investment in deep neural networks and GPUs.
- 2017Google researchers publish ‘Attention Is All You Need,’ introducing the Transformer architecture that underpins modern language models.
- 2020OpenAI releases GPT-3, demonstrating that scaling a general language model can unlock broad few-shot capabilities.
- 2022OpenAI launches ChatGPT on November 30; conversational generative AI reaches a mass audience with unusual speed.
- 2023GPT-4 arrives as governments intensify AI governance; the White House issues an executive order and the EU advances the AI Act.
- 2024The EU AI Act enters into force on August 1, establishing a phased, risk-based regulatory framework.
- 2025Stanford’s AI Index records 78% organizational AI use in 2024, $33.9 billion in generative-AI investment, and steeply falling inference costs.
Glossary
- Foundation model
- A model trained broadly enough to be adapted to many downstream tasks, rather than built for one narrowly specified job.
- Inference
- The process of running a trained model to produce an answer, prediction, image, or action; often priced by tokens or compute consumed.
- Token
- A small unit of text processed by a language model. Tokens are not identical to words, and pricing usually distinguishes input from output.
- Context window
- The amount of information a model can consider in one interaction, including instructions, retrieved documents, conversation history, and output.
- RAG
- Retrieval-augmented generation: fetching relevant external material and supplying it to a model so answers can use fresher or private information.
- Hallucination
- A fluent but unsupported or false model output. It is mitigated—not universally eliminated—through retrieval, constraints, evaluation, and review.
- Open-weight model
- A model whose trained parameters are available for download under stated terms; this does not necessarily mean its data or training process is open.
- Multimodal
- Able to process or generate more than one medium, such as text, images, audio, video, or sensor data.
- Agent
- A model-centered system that can plan steps, call software tools, inspect results, and continue toward a goal with varying autonomy.
- Benchmark saturation
- The point at which top systems cluster near a test’s ceiling, reducing that benchmark’s power to distinguish meaningful real-world capability.
FAQs
How many companies are using AI?+
Stanford’s 2025 AI Index says 78% of surveyed organizations reported AI use in 2024, up from 55% in 2023. Survey samples and definitions vary, and ‘use’ may mean anything from a pilot to a production-critical system.
Is AI getting cheaper?+
Using capable models is getting dramatically cheaper: Stanford found a greater than 280-fold decline in the cost of GPT-3.5-level inference between late 2022 and late 2024. Total product cost can still be substantial once data, integration, review, security, and compliance are included.
Is AI replacing jobs?+
Some tasks and roles will shrink, while others will expand or be redesigned. The IMF’s estimate that nearly 40% of global employment is exposed to AI measures potential interaction with work, not a forecast that 40% of jobs will disappear.
Can benchmark scores be trusted?+
They are useful under their stated conditions, especially when tests are transparent and contamination is controlled. They do not fully capture reliability, cultural fluency, latency, cost, safety, or performance inside a particular organization’s workflow.
Are open models now as good as closed models?+
Open-weight systems have narrowed the gap on several common benchmarks and can be excellent for specialized deployment. Closed frontier systems often retain advantages on the hardest tasks, while open deployment can offer control, customization, and data locality.
How large is AI’s environmental footprint?+
No single global number cleanly isolates AI from all data-center activity. The IEA puts total data-center electricity consumption at about 415 TWh in 2024 and projects roughly 945 TWh by 2030, with AI a principal growth driver.
Where is the strongest startup opportunity?+
Often above the model layer: domain-specific workflows, trusted data, evaluations, compliance, rights management, and interfaces that remove complexity. The defensible asset is usually the system of context and behavior around the model, not raw model access.
What metric should an AI product track first?+
Choose a real user outcome: successful resolutions, verified time saved, error reduction, approval rate, or revenue retained. Pair it with quality, latency, cost per completed task, and the share of outputs requiring human correction.
Predictions
- Inference prices will probably continue to fall, but premium reasoning, video generation, and agentic workloads may keep total compute spending elevated.
- Model rankings may become less strategically important as capable systems converge; proprietary context, distribution, and interaction design are likely to carry more value.
- AI agents will plausibly move from demonstrations into bounded, auditable workflows before they become trustworthy general digital employees.
- Regulation and litigation may make provenance, consent, model documentation, and rights-cleared datasets visible product features rather than back-office obligations.
- On-device and compact models are likely to grow where privacy, latency, offline operation, or predictable cost matter more than maximum frontier capability.
Risks
- Measurement theater: firms may report licenses, prompts, or pilots instead of verified improvements in quality, revenue, safety, or time.
- Provider dependence: price changes, outages, policy restrictions, and model deprecations can destabilize products built too tightly around one API.
- Synthetic abundance can degrade discovery, trust, and creative livelihoods while increasing the cost of verifying what is authentic or licensed.
- Autonomous systems can scale errors, security breaches, discrimination, and reputational damage faster than conventional human processes.
- Energy, water, chips, and grid access may become material constraints, particularly where data-center growth is geographically concentrated.
Opportunities
- Build vertical copilots around expensive, repetitive decisions where domain data and expert review can create a measurable quality advantage.
- Create evaluation and observability tools that test factuality, bias, latency, cost, and failure modes continuously in production.
- Design provenance and rights infrastructure for creative work: consent registries, attribution trails, licensing markets, and verifiable editing histories.
- Use compact or local models to make private, responsive experiences for studios, clinics, factories, classrooms, and field teams.
- Treat AI as a new design material: invent interfaces based on iteration, ambiguity, and collaboration rather than placing a chat box inside every product.
For professionals
For an operator, the most useful AI dashboard separates capability, economics, adoption, and risk. Capability should be measured on a private evaluation set reflecting actual tasks, including adversarial and low-frequency cases. Economics should track end-to-end cost per successful outcome—not token price alone—with model calls, retrieval, retries, tool usage, review, and incident handling included. Adoption should capture retained weekly use and workflow completion rather than seats provisioned. Risk should include unsupported-claim rates, data leakage, override frequency, subgroup performance, and the blast radius of autonomous actions. A model that scores lower on a public leaderboard may be the superior production choice if it is faster, cheaper, more controllable, or easier to deploy privately. Portfolio design also matters. Use a routing layer rather than assuming one model should do everything: deterministic software for rules, small models for frequent classifications, retrieval for grounded knowledge, frontier models for genuinely difficult reasoning, and humans for irreversible or accountable decisions. Preserve model portability by separating prompts, tools, schemas, memory, and evaluations from the provider interface. The resulting moat is not an API wrapper; it is accumulated workflow data, outcome feedback, institutional trust, and a product language users prefer. In a market where intelligence is increasingly abundant, disciplined context engineering and editorial judgment become forms of infrastructure.
Sources & references
- Stanford AI Index Report 2025
- International Energy Agency — Energy and AI
- International Monetary Fund — Gen-AI: Artificial Intelligence and the Future of Work
- McKinsey — The State of AI: How Organizations Are Rewiring to Capture Value
- OECD AI Policy Observatory
- European Commission — AI Act
- NIST AI Risk Management Framework
- Attention Is All You Need
| Frontier API | Adapted open-weight model | Specialized model trained in-house | |
|---|---|---|---|
| Up-front cost | Low; integration and evaluation | Medium; engineering plus hosting | Very high; data, research, compute |
| Marginal economics | Usage-based; provider sets prices | Infrastructure cost can be optimized at scale | Potentially attractive at scale, but training must be amortized |
| Control and privacy | Contract-dependent; limited model control | High when self-hosted | Highest, subject to the data supply chain |
| Time to market | Days to weeks | Weeks to months | Months to years |
| Best fit | Rapid products needing top capability | Stable, private, or customized workloads | Distinctive task with exceptional proprietary data |
| Primary risk | Vendor dependence and changing terms | Operational complexity and weaker edge-case performance | Capital intensity, obsolescence, and uncertain capability |
The Curator examines Slow Productivity for Creative Strategists through innovation scouting, tasteful design, artful technology, cultural context, product signals, future trends, and opportunity discovery, with practical signals, risks, examples, and a reason for readers to return as the story changes.
The next interface may not wait for commands. It will interpret intent, assemble tools, negotiate systems, and act—turning software from a collection of destinations into a designed field of agency.
The next era of artificial intelligence will not be decided by model scale alone. These unresolved questions—about agents, interfaces, culture, labor, trust, energy, and ownership—are where tomorrow’s products and creative possibilities are taking shape.
AI is becoming cheaper, smaller, more capable, and more culturally consequential. Here is how to read the signal beneath the benchmarks—and where builders should look next.
AI is neither a synthetic mind, an instant job-destroyer, nor an impartial machine. Seeing it clearly reveals better products, richer creative practices, and more durable opportunities.
AI leadership is no longer a single-model contest. The durable advantage belongs to those who combine capable systems with distribution, trust, distinctive data, useful interfaces, and cultural taste.