What Is General Artificial Intelligence? An Evidence Guide to AGI
The short answer
AGI is a proposed threshold for broad machine intelligence—not a synonym for a chatbot, agent or general-purpose model.
Definition first
Different organisations set different thresholds. A claim is meaningless until it names the definition.
Evidence across dimensions
Breadth, depth, transfer, reliability and test conditions matter more than a single headline score.
Act on present systems
Australian teams still need use-case governance, testing and human control whether or not a vendor uses the AGI label.

Direct answer
AGI means broad capability—but the threshold is disputed.
Artificial general intelligence usually describes a proposed AI system with strong capability across a wide range of cognitive tasks, including unfamiliar ones. There is no single operational definition accepted by every lab, researcher and government. That is why “this is AGI” is a claim to evaluate, not a self-proving product category.
A terminology note matters here: AGI conventionally expands to artificial general intelligence. “General artificial intelligence” is a common word-order variation and the search phrase used in this page’s original title. Both usually point to the same debate.
Four terms that should not be collapsed into one
The clearest way through the hype is to separate the category of system, the breadth of its uses, the proposed capability threshold and the way it operates.
| Term | Source or lens | What it means here | What the label does not prove |
|---|---|---|---|
| AI system | OECD policy definition | A machine-based system that infers how to produce outputs such as predictions, content, recommendations or decisions. | Human-level ability, broad competence or autonomy. |
| General-purpose AI | International AI Safety Report 2026 | Models and systems able to perform a wide variety of tasks. | That performance is human-level, reliable across every domain or AGI. |
| AGI | OpenAI Charter | OpenAI’s threshold refers to highly autonomous systems outperforming humans at most economically valuable work. | Industry-wide agreement; it is one influential organisation’s definition. |
| Levels of AGI | Google DeepMind research framework | A proposed classification using performance depth and capability breadth, considered separately from autonomy and deployment risk. | That a single benchmark result fixes a system’s level across all tasks. |
| AI agent | System-design lens | A system that can select and perform steps toward a goal with some degree of independence. | Generality: a narrow agent may still act autonomously inside one workflow. |
These definitions answer different questions. The OECD definition helps policy makers decide what counts as an AI system. The 2026 International AI Safety Report discusses the capabilities and risks of today’s most capable general-purpose systems. OpenAI’s Charter states a mission-specific AGI threshold. DeepMind’s framework tries to make progress more measurable by separating breadth from performance. None should be silently substituted for another.
What evidence would make an AGI claim meaningful?
Start with breadth: which materially different tasks and domains are represented? A model completing code, prose and questions through the same text interface may be broadly useful, but the test still needs to show what population of human work those tasks represent.
Then ask about depth: compared with whom, at what performance level and with what reliability? “Human-level” could mean an untrained participant, a typical worker or a domain expert. Average performance can also hide catastrophic failures that matter in a real deployment.
Finally, look for transfer and adaptation. An important test uses unfamiliar, held-out problems and discloses what training, prompting, tools, examples and human intervention were allowed. Otherwise a result may measure preparation for a test rather than the ability to acquire a new skill.
MLAI evidence worksheet
Stress-test an AGI claim
Paste or summarise the claim for your own working notes, then check only what its published evidence supports. Nothing is submitted or saved.
Evidence score 0 / 8
Poorly specified
There is not enough information to assess the claim. Ask for a definition, test protocol and complete results.
Why one benchmark cannot settle the question
A benchmark is an instrument, not a verdict. It samples particular abilities using a particular interface, dataset, scoring rule and budget. Strong performance can be genuine progress while still leaving other definitions of AGI unresolved.
ARC-AGI is a useful example because its designers deliberately focus on skill acquisition and novel problems rather than accumulated knowledge. The 2026 ARC-AGI-3 technical report introduces interactive environments that require exploration, goal inference and planning. It reports that humans in its study solved all tested environments while evaluated frontier systems scored below one per cent as of March 2026. That is evidence about those systems on that benchmark under its protocol—not a universal measurement of every component of intelligence.
Minimum questions for any benchmark headline
- Was the task set held out, and how was possible training-data exposure investigated?
- Which prompts, tools, memory, retries, time and human help were allowed?
- Is the reported result an average, best run or selected demonstration?
- What do humans score under comparable time, information and interface conditions?
- Which failure modes, costs and safety constraints disappear from the headline number?
- Can independent evaluators reproduce the protocol and challenge the conclusion?
What can be said responsibly in July 2026?
The International AI Safety Report 2026 says general-purpose AI capabilities have continued to improve, including through techniques applied after initial training. It also describes several plausible paths to 2030: progress could slow, continue at recent rates or accelerate. That range is a useful antidote to treating one forecast as settled fact.
This guide therefore does not declare that a particular product is or is not AGI for every possible definition. It makes a narrower claim: no shared label removes the need to define the threshold, publish the test conditions and examine contrary evidence. An organisation may announce that its own threshold has been reached while independent researchers reasonably reject the definition or evidence.
Responsible statement
“Under definition D and protocol P, the system exceeded baseline B on these task groups; these failures and open questions remain.”
Weak statement
“It reasoned in our demo, so AGI is here.”
What this means for an Australian builder today
A team does not need an AGI forecast to make a useful AI decision. Start with a real workflow, accountable owner, permitted data, measurable outcome and human intervention point. Test the system that will actually be deployed, including predictable failures, rather than borrowing a frontier-model headline.
The Australian Government National AI Centre’s current foundations guidance recommends accountability, risk management, transparency and human control appropriate to the use. Those controls follow from what the system does and whom it affects—not from whether marketing calls it AI, an agent or AGI.
- Learn the basic distinction first in MLAI’s plain-English guide to AI.
- If the system takes actions, map its boundaries using the guide to AI agents.
- Bring disputed claims and methods to an MLAI event for discussion with the Australian community.
- When you have a defined workflow and success measure, scope the practical system through MLAI Studio.
Primary definitions and evidence
Sources checked for this 28 July 2026 revision.
[1]Levels of AGI for Operationalizing Progress on the Path to AGI
Google DeepMind, ICML 2024 • Research framework separating breadth or generality, performance depth and deployment factors such as autonomy and risk.
Analysis[2]OpenAI Charter
OpenAI • An organisation-specific AGI definition centred on highly autonomous systems and economically valuable work.
Industry[3]International AI Safety Report 2026: Executive Summary
International AI Safety Report • Independent expert synthesis on current general-purpose AI capabilities, risks and uncertainty about progress.
Analysis[4]Explanatory memorandum on the updated definition of an AI system
OECD.AI • Policy definition of an AI system, useful for separating the broad category of AI from claims about AGI.
Government
Show all 6 references (2 more)Show less
[5]ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence
ARC Prize Foundation, 2026 • Technical report for an interactive benchmark focused on adaptation to novel environments, including its stated scope and results.
Analysis[6]Guidance for AI adoption: foundations
Australian Government National AI Centre • Current Australian guidance on accountability, risk, transparency, testing and appropriate human control.
Government
AGI questions, answered carefully
What does AGI stand for?
Is general-purpose AI the same as AGI?
Has AGI been achieved?
Does passing one benchmark prove AGI?
Is an autonomous AI agent necessarily AGI?
What should Australian teams do about AGI now?
Disclaimer: This article provides general information and is not legal or technical advice. For official guidelines on the safe and responsible use of AI, please refer to the Australian Government’s Guidance for AI Adoption →
Join our upcoming events
Connect with the AI & ML community at our next gatherings.
Check availability, any waitlist and registration requirements on the organiser’s page. A listing does not reserve a place.
VICTOR:AI Startups + AI Networking Night
View event details (opens in a new tab)
Hack Your Way To Page #1 on Search Engines with AI
View event details (opens in a new tab)
AI Engineer Night Melbourne #2
View event details (opens in a new tab)
Need an AI workflow, harness or agent built?
Tell MLAI Studio what outcome you need, what systems are involved and how success should be measured.
