Disclaimer: This article provides general information and is not legal or technical advice. For official guidelines on the safe and responsible use of AI, please refer to the Australian Government’s Guidance for AI Adoption →
What Is General Artificial Intelligence? An Evidence Guide to AGI
The short answer
AGI is a proposed threshold for broad machine intelligence—not a synonym for a chatbot, agent or general-purpose model.
Definition first
Different organisations set different thresholds. A claim is meaningless until it names the definition.
Evidence across dimensions
Breadth, depth, transfer, reliability and test conditions matter more than a single headline score.
Act on present systems
Australian teams still need use-case governance, testing and human control whether or not a vendor uses the AGI label.
Direct answer
AGI means broad capability—but the threshold is disputed.
Artificial general intelligence usually describes a proposed AI system with strong capability across a wide range of cognitive tasks, including unfamiliar ones. There is no single operational definition accepted by every lab, researcher and government. That is why “this is AGI” is a claim to evaluate, not a self-proving product category.
A terminology note matters here: AGI conventionally expands to artificial general intelligence. “General artificial intelligence” is a common word-order variation and the search phrase used in this page’s original title. Both usually point to the same debate.
Four terms that should not be collapsed into one
The clearest way through the hype is to separate the category of system, the breadth of its uses, the proposed capability threshold and the way it operates.
Term
Source or lens
What it means here
What the label does not prove
AI system
OECD policy definition
A machine-based system that infers how to produce outputs such as predictions, content, recommendations or decisions.
Human-level ability, broad competence or autonomy.
General-purpose AI
International AI Safety Report 2026
Models and systems able to perform a wide variety of tasks.
That performance is human-level, reliable across every domain or AGI.
AGI
OpenAI Charter
OpenAI’s threshold refers to highly autonomous systems outperforming humans at most economically valuable work.
Industry-wide agreement; it is one influential organisation’s definition.
Levels of AGI
Google DeepMind research framework
A proposed classification using performance depth and capability breadth, considered separately from autonomy and deployment risk.
That a single benchmark result fixes a system’s level across all tasks.
AI agent
System-design lens
A system that can select and perform steps toward a goal with some degree of independence.
Generality: a narrow agent may still act autonomously inside one workflow.
These definitions answer different questions. The OECD definition helps policy makers decide what counts as an AI system. The 2026 International AI Safety Report discusses the capabilities and risks of today’s most capable general-purpose systems. OpenAI’s Charter states a mission-specific AGI threshold. DeepMind’s framework tries to make progress more measurable by separating breadth from performance. None should be silently substituted for another.
What evidence would make an AGI claim meaningful?
Start with breadth: which materially different tasks and domains are represented? A model completing code, prose and questions through the same text interface may be broadly useful, but the test still needs to show what population of human work those tasks represent.
Then ask about depth: compared with whom, at what performance level and with what reliability? “Human-level” could mean an untrained participant, a typical worker or a domain expert. Average performance can also hide catastrophic failures that matter in a real deployment.
Finally, look for transfer and adaptation. An important test uses unfamiliar, held-out problems and discloses what training, prompting, tools, examples and human intervention were allowed. Otherwise a result may measure preparation for a test rather than the ability to acquire a new skill.
MLAI evidence worksheet
Stress-test an AGI claim
Paste or summarise the claim for your own working notes, then check only what its published evidence supports. Nothing is submitted or saved.
Evidence score 0 / 8
Poorly specified
There is not enough information to assess the claim. Ask for a definition, test protocol and complete results.
Why one benchmark cannot settle the question
A benchmark is an instrument, not a verdict. It samples particular abilities using a particular interface, dataset, scoring rule and budget. Strong performance can be genuine progress while still leaving other definitions of AGI unresolved.
ARC-AGI is a useful example because its designers deliberately focus on skill acquisition and novel problems rather than accumulated knowledge. The 2026 ARC-AGI-3 technical report introduces interactive environments that require exploration, goal inference and planning. It reports that humans in its study solved all tested environments while evaluated frontier systems scored below one per cent as of March 2026. That is evidence about those systems on that benchmark under its protocol—not a universal measurement of every component of intelligence.
Minimum questions for any benchmark headline
Was the task set held out, and how was possible training-data exposure investigated?
Which prompts, tools, memory, retries, time and human help were allowed?
Is the reported result an average, best run or selected demonstration?
What do humans score under comparable time, information and interface conditions?
Which failure modes, costs and safety constraints disappear from the headline number?
Can independent evaluators reproduce the protocol and challenge the conclusion?
What can be said responsibly in July 2026?
The International AI Safety Report 2026 says general-purpose AI capabilities have continued to improve, including through techniques applied after initial training. It also describes several plausible paths to 2030: progress could slow, continue at recent rates or accelerate. That range is a useful antidote to treating one forecast as settled fact.
This guide therefore does not declare that a particular product is or is not AGI for every possible definition. It makes a narrower claim: no shared label removes the need to define the threshold, publish the test conditions and examine contrary evidence. An organisation may announce that its own threshold has been reached while independent researchers reasonably reject the definition or evidence.
Responsible statement
“Under definition D and protocol P, the system exceeded baseline B on these task groups; these failures and open questions remain.”
Weak statement
“It reasoned in our demo, so AGI is here.”
What this means for an Australian builder today
A team does not need an AGI forecast to make a useful AI decision. Start with a real workflow, accountable owner, permitted data, measurable outcome and human intervention point. Test the system that will actually be deployed, including predictable failures, rather than borrowing a frontier-model headline.
The Australian Government National AI Centre’s current foundations guidance recommends accountability, risk management, transparency and human control appropriate to the use. Those controls follow from what the system does and whom it affects—not from whether marketing calls it AI, an agent or AGI.
ARC Prize Foundation, 2026 • Technical report for an interactive benchmark focused on adaptation to novel environments, including its stated scope and results.
Australian Government National AI Centre • Current Australian guidance on accountability, risk, transparency, testing and appropriate human control.
Government
AGI questions, answered carefully
What does AGI stand for?
AGI usually stands for artificial general intelligence. “General artificial intelligence” is a common word-order variation, but artificial general intelligence is the standard expansion used by the sources in this guide.
Is general-purpose AI the same as AGI?
No. General-purpose AI can perform a wide variety of tasks, but that does not establish human-level breadth, transfer, reliability or autonomy. The International AI Safety Report uses general-purpose AI for present systems without treating the term as proof of AGI.
Has AGI been achieved?
There is no shared operational definition or single external test that settles that question for every researcher and organisation. A responsible claim must name its definition and evidence instead of presenting a disputed label as a universal fact.
Does passing one benchmark prove AGI?
No. A benchmark samples particular capabilities under particular conditions. Useful evidence also reports task breadth, novelty, baselines, repeated trials, failures, assistance, contamination risks and independent replication.
Is an autonomous AI agent necessarily AGI?
No. Autonomy describes how independently a system pursues goals or performs actions. Generality describes the breadth of capability. A narrowly capable system can be highly autonomous, while a broad model can still require frequent human direction.
What should Australian teams do about AGI now?
Make decisions using the capabilities and risks of the system being deployed today. Define the use case, test it, assign accountability, keep appropriate human control and document limitations rather than relying on an AGI label or forecast.