Wisma
Method11 May 2026· Wisma

What Your Research Agent Doesn't Tell You

A version of this comes up in nearly every recent ministry conversation we have: “We use [Deep Research / OpenClaw / Hermes Agent] for this, and the output seems fair — it just doesn't quite hit the mark. What can we do to improve it?”

It is a good question, asked by teams doing the sensible thing. They picked a capable tool, ran their research through it, and got back a memo that reads well and cites its sources. Nothing in it looks wrong, but for some reason it just doesn't hit the mark, and the instinct is to reach for a better prompt.

Sometimes that is the answer. Often it is not, because the shortfall is not in how the question was asked. It is in what never got read. An agent that cannot see a source does not report a gap; it writes around it. An agent that cannot read a language reports on the part it could read. What comes back is a confident, complete-looking document with an unstated boundary around what went into it. No prompt, however carefully phrased, moves that boundary.

WHAT THE RUN SURFACEDTHE BOUNDARY A PARTNER DRAWS
Every source, thread, and voice in the field — only the lit cluster was reached.

Start with what these tools do well. A research agent will browse on its own for minutes at a stretch, read text, images, and PDFs, follow leads it turns up mid-run, and return a long, cited synthesis. Many now let you name or favour particular sites, and connect to outside data through MCP. Some drive a real browser, which takes them past what a search index alone would surface. For a one-shot literature review on an indexed topic — a sector primer, an industry scan, a precedent search — this is often the right tool. We use these tools ourselves for that work, and they give a quick, useful first pass.

The gap opens when the brief needs more than a literature review. Sometimes the brief is find and synthesise what has been said about X in this period. Sometimes it is tell me how citizens took [policy announcement] in the 24 hours after, on the platforms where they actually argue, in the languages they actually speak, separated from coordinated noise, with confidence levels we can stand behind.

The first is research. The second is intelligence.

Both are real jobs, and they need different tools. This piece does not argue that research agents are bad — they are not — and it does not run a feature table. It offers a way to tell which job is in front of you, and to see the boundary the agent's own output never draws.

A worked example: the post-event reception read

Imagine a high-stakes policy announcement made six hours ago. Senior management wants an hourly read for the first 24 hours, until things settle, then a full report: how did it land among the grassroots, who spoke up and on which platforms, are early narratives forming that will need an answer, and is the reaction growing on its own or being amplified.

What a research agent returns. A polished memo. Mainstream coverage well summarised. A handful of high-engagement X posts drawn into the synthesis. Surface themes named, with citations. The prose reads confident and well formed. For a quick read it is genuinely useful.

What it misses on this brief. Six things, each pointing to one of the gaps — and none of them visible in the memo.

  • The HardwareZone EDMW thread where the critique first seeded, and where the most-shared screenshots came from, never appears. The agent did not reach it on this run. Even when it does pull in one representative Reddit thread, the rest of social media stays out. (Gap 1: source control.)
  • The first hours' reactions are missing, because web indexing has not caught up with them. Meanwhile material from outside the 24-hour window comes in because it ranks well in search, watering down the post-event signal. (Gap 2: bounded time.)
  • Mainstream articles are read article by article, but the social posts and the comment threads beneath them — where the reception is actually argued out — are never read at all. What the agent worked from is a sample of citations, not the conversation. (Gap 3: depth of corpus.)
  • Only English is covered. Mandarin, Malay, Tamil, and code-switched Singlish are not, and the meaning carried in them is lost. (Gap 4: native multilingual.)
  • Comments from accounts created days ago, which may be bots or spam, are counted as authentic public reaction alongside organic voices in the “what people are saying” tally. (Gap 5: authentic vs amplified.)
  • A confident claim of a shift in sentiment that the underlying data does not support. Trying to be helpful, the model writes fluent prose over thin evidence. (Gap 6: calibrated claims.)

Read the memo on its own and none of these register as absences. It does not say I could not reach the forums, or I read this only in English, or my window leaked. That is the point. The boundary is invisible from the output, so it never comes up when you decide how far to trust the read.

What a Wisma post-event report does on the same brief. We set the sources first, across mainstream, social, and forums, including closed spaces where lawful and appropriate. We hold the scope to the actual 24-hour window, not to “top results.” We read the conversation in native English, Mandarin, Malay, and Tamil, with English translations for analyst checks, keeping the cultural and specialist sense intact. We strip bots and spam out of the dataset before the analysis runs, so it can tell organic reception from inauthentic amplification. And we record source, method, and confidence on every classification, with a human reviewer checking the result.

Six gaps to close

Each gap in turn.

Six gaps between what a research agent returns and what a brief needs.

1. Source control, not just search. What a research agent reads is whatever its tools surface on that run. Site-priority settings let you name or favour indexed sites, which helps when your sources happen to be indexed, and an agent driving a real browser reaches further than a search index alone. But reach is not coverage: such an agent can open a forum page; it cannot tell you how much of that forum's discussion it read. The material also moves with the phrasing: petrol prices, petrol costs, and oil prices turn up different sources and can support different conclusions, and the memo shows you the run you got, never the spread across the runs you did not do. And the open web is closing. More and more legitimate sites now block AI crawlers or throttle automated access, so the reachable material thins even as the agents get better at reading it. What no agent does is weigh sources by trust and method beyond an on/off toggle, or account for what it never reached. A media intelligence partner sets the sources on the question's terms — which forums, which platforms, which closed spaces, each weighted by trust and by relevance to the brief — and says what is in scope, so you know what the answer rests on.

2. Time, bounded to the question — the very recent and the very specific. Two common failure modes: (a) Sub-day windows. When the question covers the first twelve hours after an announcement, the agent waits on search indexing, which for many sources runs into the next day or later. A partner ingesting news, social, and forums continuously sees the same material within minutes of publication. (b) Historically bounded windows. When the question is “13–20 January 2024,” the agent returns the best relevance matches for the search terms, and those routinely include material from outside the range. Ask about discussions in April and you get the wrong year, articles that mention April in passing, and the odd item where April is somebody's name. Relevance ranking is not a date filter, and the memo will not mention the difference. A partner takes the start, the end, and the granularity as part of the question, and reads only what falls inside.

3. Depth of corpus, not a sample of it. Research agents work page by page. They pick high-relevance sources and synthesise from a finite set of citations, and they are good at it. But when the answer requires what most people are saying — every comment thread under the article, every forum reply, every related chain on social — the unit changes from cited sources to data points, and the volume rises by orders of magnitude. However many sources it cites, no single-context-window run works at that volume. Limited context and compute force it to pick a subset, and the sampling does not show in the output. What a sample loses is specific: minority views, narratives still forming, and the true weight of a sentiment — how much of the conversation holds it, as against how loudly it appears. A partner reads the conversation at its full depth and filters down to what matters, so the read rests on the conversation itself rather than a sample of citations from it.

4. Native multilingual, with cultural and specialist context. Singlish, dialect phrases, code-switched threads, in-group sarcasm: all carry meaning that machine translation flattens. How much survives depends on the model underneath and on the quality of the translation layer, and the reader of the memo cannot see when it fails. Models that natively understand multilingual and multicultural contexts read this material in the register it was written in. Specialist fields compound the problem — legal language, healthcare terms, technical trade idiom each have conventions a translation pass cannot rebuild. There is a quieter version of this failure too. Prompted in English, an agent tends to search in English, then reports what it found without noting that the other languages were never queried. A partner builds for that linguistic and cultural ground deliberately, rather than treating English translation as the default unit of analysis.

5. Authentic signal, not amplified noise. Bot rings, coordinated reply networks, persona farms, and brigading inflate engagement counts and skew the “what people are saying” tally. A partner strips these from the dataset before the analysis runs, so the read describes organic reception rather than amplification machinery. The honest, workable form of this sits at the network layer, not the content layer: what amplified this, who seeded it, which network carried it. Whether any single post was written by a human is not something we, or anyone, can reliably tell today — and a partner who claims otherwise should not be trusted on the rest of the stack either.

6. Calibrated claims, not confident prose. The model providers say this themselves. OpenAI's documentation for ChatGPT Deep Research grants that the tool “occasionally makes factual hallucinations or incorrect inferences” and “may not accurately convey uncertainty,” and the same caveat, in other words, applies across the class. The disclosure is worth holding onto, because it names the real risk: not that the agent is often wrong, but that its confidence does not move with its evidence. There is a second-order version to watch for. The model is trying to help. It fills gaps in thin data, and it leans towards material that seems to support what it takes you to be looking for. A brief carrying an expectation can come back confirmed, fluently and with citations. What matters to the buyer is what method sits on top of the model. A partner pairs it with proprietary hallucination-reduction techniques and human review on classifications that carry political weight, and records source, method, and confidence on every published claim. The result is something you can still cite in a week, in another register, with the method notes intact.

Research, intelligence, and what to do with the difference

The six gaps do not add up to a verdict on which tool is better. They mark the line between research and intelligence — between find and synthesise what has been written about X and tell me what is happening in this material, on this timeline, in these languages, separated from this noise, with confidence levels you can stand behind. Most ministry briefs sit cleanly on one side or the other. Some sit on the research side, and a research agent is the right tool for them. Some sit on the intelligence side, and that needs a different kind of partner.

ResearchOne-shot

“Find and synthesise what has been said about X in this period.”

A literature review on an indexed topic — a sector primer, an industry scan, a precedent search. For this, a research agent is often the right tool.

IntelligenceStanding

“Tell me how citizens took this announcement in the 24 hours after — on the platforms where they argue, in the languages they speak, separated from coordinated noise, with confidence we can stand behind.”

A different job — and a different kind of partner.

Two briefs that look alike — and need different tools.

The difficulty is that both jobs produce documents that look alike. Both come back polished, cited, and confident. You cannot tell them apart by reading the output more closely, or by rewriting the prompt. You tell them apart beforehand, by asking what the brief actually needs: which sources must be in scope, which window, which languages, and what follows if part of the conversation is missing. When any of those answers bears on the decision, an unstated boundary is not something you can absorb.

We should say openly that we are one of the companies building to close these gaps. NarrativeIQ is built for the intelligence job this piece describes:

  • Sources set in advance across mainstream, social, and forums, including closed spaces where lawful and appropriate.
  • Continuous ingestion at minute-level latency, with the analysis held to the window you set and read at the depth of the conversation, not a sample of it.
  • Models that natively understand multilingual and multicultural contexts, including specialist and in-group terms.
  • Spam and coordinated amplification stripped from the dataset before the analysis runs.
  • Source, method, and confidence recorded in our reports, with domain reviewers checking the claims that carry weight.

If you are working out how research becomes intelligence in your own team, and how to tell brief by brief which one you need, we would be glad to talk it through and share what we have seen across similar work. Sometimes the honest answer is use a research agent. Other times it is not. Write to us at contact@wisma.ai.