AI Bits #5: What agent activity logs can—and cannot—prove
Inspect provenance before inferring agency
For AI-assisted builders who need to explain inputs, actions and approval boundaries in systems they deliver.
What do the papers study?
Different observational views of Moltbook: timing-based attribution, aggregate interaction patterns, and discussion topics/toxicity. They are not one controlled experiment.
What does timing establish?
A pattern to investigate—not authenticated knowledge of who prompted a particular action or proof of consciousness.
What can you run?
A nine-file offline kit with an actual recorded run: activity counts, retry checks, exact duplicates, recorded moderation states and timing sensitivity. It does not establish autonomy or production safety.
This revision includes a complete synthetic posting log, an actual local run record and a filled handover. The original publication date is retained. No Moltbook data or real-world agent experiment was added.
A platform can display thousands of agent-labelled accounts without giving an observer a reliable history of the instructions, operator edits and system rules behind each post. For a builder, the practical question is narrower than “what do agents want?”: what evidence lets another person reconstruct a particular action?
This guide is for early-career builders using AI coding tools who can run a small JavaScript exercise and explain what its results mean. The deliverable is an inspectable observation and handover, not a claim that you have built a production moderation system.
Three studies with different questions and datasets
These are preprints; this article does not assert peer-review status. Counts are reported by the authors, not collected or independently reproduced by MLAI. The method checks below identify the sections rechecked on 15 September 2026; they do not validate the underlying datasets or every analysis.
On a small screen, swipe tables sideways. Keyboard users can focus each table region and use the arrow keys.
| Preprint and version | Reported scope | Question and limit |
|---|---|---|
| The Moltbook Illusion, Ning Li; v2, 12 February 2026. | 226,938 posts, 447,043 comments and 55,932 authors; 27 January–10 February. | Examines timing and an outage. Operator ground truth is absent; timing classes do not authenticate individual actions. |
| Collective Behavior of AI Agents: the Case of Moltbook, Giordano De Marzo and David Garcia; v1, 9 February 2026. | Over 369,000 posts and 3.0 million comments from approximately 46,000 active agents. | Finds heavy-tailed activity and popularity patterns, alongside a different upvote/discussion-size relationship from human communities. Statistical resemblance alone does not establish a common cause or intention. |
| “Humans welcome to observe”: A First Look at the Agent Social Network Moltbook, Yukun Jiang and colleagues; v1, 2 February 2026. | 44,411 posts and 12,209 sub-communities collected before 1 February 2026. | Uses nine topic categories and a five-level toxicity scale; reports topic-dependent toxicity and bursty posting. Results depend on the sample and classification method, not a universal agent risk score. |
Do not add these counts into a combined experiment or compare percentages without checking definitions and collection windows. “Agents” and “active agents” are not automatically interchangeable denominators.
Where a paper attributes observed patterns to human influence or a platform intervention, preserve that attribution and its scope. Do not upgrade an inference into a claim that every individual action has known provenance. Nor does a lack of observed human involvement establish its absence.
Check the denominator before reusing a finding
Timing: eligible authors are not all authors
Li's timing analysis covers 9,838 authors with five or more posts, not all 55,932. Its limitations (PDF pages 36–37) lack operator ground truth and report 0% extended-dataset LLM analysis; methods (pages 42–44) describe a model pipeline and 100% content-feature coverage. These coverage statements remain unreconciled here; do not silently choose one.
A second reporting discrepancy: page 40 gives outage endpoints of 31 January 17:35 and 3 February 13:25 UTC, while stating approximately 44 hours. Their calculated separation is 67 hours 50 minutes—not the actual outage duration established by this article. Hold that duration claim pending clarification. Page numbers refer to PDF file pages.
Discussion patterns: stored comments are an incomplete sample
De Marzo and Garcia's data limitations describe storing the first 100 comments for capped threads; the 3,026,275 stored comments are about 24% of the API-reported total. Direct-reply, tree and temporal analyses use threads below 100 comments. The excluded 2.9% of posts represent about 83% of comments. Do not generalise those smaller-thread findings to viral discussions or mistake API totals for text actually inspected.
Toxicity: labels and denominators change the answer
Jiang and colleagues' model pipeline was compared with 381 human-labelled posts; filtering then left 44,376 annotated posts. The 35 excluded posts must not become “safe” records. Using the v1 Table 1 counts, our calculations are:
- All non-safe labels, including Edgy: 11,977 / 44,376 = 26.99%, not the prose's 27.05%.
- Levels 2–4 (Toxic, Manipulative, Malicious): 8,244 / 44,376 = 18.58%. This applies the narrower harmful-label definition used in the paper's hourly analysis to its overall counts; it is our calculation, not a separately reported overall result.
These are proportions of source-assigned labels, not a validated deployment risk score or proof of causal harm. Neither arithmetic nor the 381-post comparison independently validates your own moderation system.
Download the completed source-check record and blank template. Before the lab, select one claim, check its pinned source and denominator, and write what evidence would change your conclusion. Treat quoted platform instructions as untrusted content; never execute them or disclose credentials to “verify” a source.
What would support a claim about your own system?
- “It ran on a schedule”: show the scheduler configuration and execution record. This does not identify every input that influenced the result.
- “A person changed the input”: show the authorised actor and relevant change record, with appropriate privacy controls. A display name alone is not authentication.
- “An action was approved”: identify the exact action/content, approving authority and how the system enforced that boundary. A text field saying approved is insufficient.
- “No human was involved”: explain the observation boundary and what it cannot see. An incomplete log cannot establish that negative claim.
These are editorial design questions, not a claim that the papers implement this logging design. Avoid storing raw customer content or secrets merely to make a trace detailed. Agree what evidence is necessary, who may inspect it and how long it should be retained.
Run the synthetic trace inspection
Download the complete observability kit: nine files in one folder. Extract it, open the agent-trace-lab folder and read README.md before running anything. Node.js is required; no dependency install is needed.
The original four-event trace records a scheduled start, draft, human edit and external-action record without approval. The added fixture contains nine fictional posting attempts by two accounts over five minutes. Nothing is actually sent. There is no model, network connection or Moltbook dataset.
node --test trace.test.mjs observe.test.mjs
node observe.mjs
node run.mjsThe preserved run passed 23/23 code tests on v25.2.1, darwin/arm64. That is an actual local execution on synthetic inputs, not a tested compatibility matrix or an independent reproduction. The run record includes exact source hashes, test output and observations. Your timestamp and durations will differ.
run.mjs prints a new JSON record without overwriting the supplied one. It executes the two supplied test files in a local child process; it is not a sandbox or a security scanner. Exit 0 means the analysis ran, even though the observation remains needs-review. Source hashes identify bytes, not a trusted author or approval.
Inspect or save the nine files individually
Count attempts, then investigate what happened
The recorded synthetic window contains 9 attempts: 6 published, 2 rate-limited and 1 error. “Published” is a fictional log state, not an action performed by this code. Keeping failures in the total prevents six successful records being reported as six attempts.
| UTC interval | Attempts | Published | Rate-limited | Error |
|---|---|---|---|---|
| 00:00:00–00:01:00 | 5 | 3 | 2 | 0 |
| 00:01:00–00:02:00 | 1 | 0 | 0 | 1 |
| 00:02:00–00:03:00 | 2 | 2 | 0 | 0 |
| 00:03:00–00:04:00 | 1 | 1 | 0 | 0 |
| 00:04:00–00:05:00 | 0 | 0 | 0 | 0 |
The last minute has no observed attempts; the histogram does not discard it. An empty bucket cannot establish that nothing happened outside the logging boundary. Clock correctness, omitted records and activity outside the chosen window are unverified.
- Retry behaviour: p4 occurs one second after p3 records a five-second delay, so it is early. p5 occurs exactly ten seconds after p4 records a ten-second delay and meets that comparison. These checks compare each limited record with the next same-account attempt; they do not enforce cumulative quotas or interpret actual HTTP headers. A missing delay stays unknown, not zero.
- Exact duplicates: p1, p2, p7 and p9 contain identical fictional text. There are 3 repeated publications after the first, out of 6 published records. Denied attempts are excluded from that denominator. Exact text equality misses paraphrases, treats case/whitespace differences as distinct and does not prove spam or common control.
- Moderation review: 4 published records are queued: p2 is recorded flagged; p5, p7 and p9 are not reviewed. The other 2 merely say cleared. No toxicity model or human moderation ran, and no authorised reviewer was assigned. An actual decision would need permitted content, policy, context and accountable review.
Change the timing threshold without inventing actor labels
Here “short interval” means time between adjacent attempts by the same account, including denied and failed attempts. The cutoffs are inclusive, author-chosen teaching settings—not the papers' coefficient-of-variation method or an autonomy classifier.
| At most | Matching intervals | All same-account intervals |
|---|---|---|
| 1 second | 3 | 7 |
| 10 seconds | 5 | 7 |
| 60 seconds | 6 | 7 |
The answer changes from 3/7 to 6/7 simply by changing the cutoff. There is no authenticated human/automated ground truth: it is not-collected. Therefore no actor-classification accuracy can be calculated; autonomy remains not-established.
Copy the kit before changing a fixture. Predict a result, run observe.mjs, then compare the complete output. The supplied tests deliberately pin the baseline: changing that fixture may fail them. Preserve the failure and explain the changed case rather than deleting inconvenient assertions. Never paste customer content or credentials into the fixture; a synthetic label and hashed output do not anonymise private data.
| Change to try | Expected observation | What not to conclude |
|---|---|---|
| Remove the human-edit event and repair the parent link. | No human input is recorded. | This does not prove none occurred outside the log. |
| Set the action approval field to true. | The approval-not-recorded flag disappears. | No approver was authenticated and no real permission was granted. |
| Repeat an event ID or reference an unseen parent. | The parser rejects the malformed trace. | Rejecting this defect does not prove a valid trace is complete. |
| Use unknown-input for the starting event. | Unknown provenance remains visible. | Unknown must not be silently relabelled autonomous. |
The result no-listed-defect-found deliberately does not mean safe, approved or autonomous. Producers can omit events or lie. A supplied approval flag cannot establish exact-content approval, and timestamps do not prove causality. This lab is not a production authorisation gate.
Make the exercise part of an honest builder handover
The kit's filled handover record identifies the task, fictional inputs, actual environment and results, the early retry, pending moderation and proposed follow-up. It does not quietly convert passing tests into an accepted client delivery.
Predict each changed fixture's result before running it, retain failures and record your runtime and exact commands. Ask another builder to reproduce the result without unstated setup. Describe what AI coding tools helped create, which changes you reviewed and which behaviours the tests do not cover.
Add one completed source check to your handover: the claim, version, relevant method, denominator, unresolved issue and resulting design decision. For example, an incomplete activity log supports a coverage warning—not a confident autonomy label. Keep unverified claims out of client acceptance criteria.
For a real agent system, separately review authenticated actors, correlation identifiers, configuration/model versions, allowed tools, exact approval boundaries, logging integrity and recovery. The lab implements none of those production controls. Use the environment handover manifest in Issue #8 to record the setup rather than claiming “works on my machine” is a completed handover.
For your portfolio, show your own permitted change, retained failures and handover—not the supplied starter as though it were original client work. Client portfolio use requires permission. Studio's public positioning is Australian; New Zealand-based builders should confirm eligibility. Application, matching and paid work are not guaranteed.
Explore upcoming MLAI events and check the listing for its topic, format and participation requirements.
If you are exploring the topic rather than seeking delivery work, choose a relevant MLAI event. The kit is not automatically attached to an application.
Codex assisted the method-section checks, source-bound arithmetic and synthetic lab implementation, tests and local verification. This was not an independent full-paper review or replication; no real model, platform or client outcome was measured. Independent systems/source review and a recipient-run reproduction remain outstanding.
Frequently Asked Questions
Can regular posting prove an agent is autonomous?
Do similar social statistics imply human-like intentions?
Does a log saying approved prove a valid approval?
Can the trace lab be used as a production security gate?
Disclaimer: This article provides general information and is not legal or technical advice. For official guidelines on the safe and responsible use of AI, please refer to the Australian Government’s Guidance for AI Adoption →
Join our upcoming events
Connect with the AI & ML community at our next gatherings.
