Every open-source investigation eventually meets the same question, and it is the one that decides whether the work was worth anything: how do you know?

Anyone with a browser can find something. The hard part, the part that separates an investigator from a search-engine operator, is turning what you found into a claim you can defend when a hostile attorney, a skeptical editor, or a decision-maker with real consequences on the line pushes back.

Investigations tend to fail on that question in the same few ways. Someone matches an image to a location instead of trying to disprove it, and matches the wrong one. Someone finds a decisive post and screenshots it a week later, after it has been edited, with nothing preserved. The finding was real. The proof was not.

OSINT gets taught as a tour of tools and websites. That is backwards. Tools change every quarter. Platforms lock down, sites die, and this year's clever trick is next year's dead link. What lasts is judgment: framing a question, planning collection, weighing sources, reasoning under uncertainty, guarding against bias, and writing findings that survive scrutiny.

What follows is that discipline in eight practices, plus the one section most field guides skip. Every example is generic. The methods are not.

1. Frame the mission before you open a browser

Most bad investigations are lost before collection starts, because someone accepted a vague tasking and never turned it into an answerable question. "Find everything on this company" is not a question. No boundary, no success criterion, no stop condition. It produces a pile of material and no judgment.

Sort what you have into three buckets first:

Then define your Priority Intelligence Requirements, the two or three decision-driving questions your customer actually needs answered, and under each one the Essential Elements of Information, the concrete sub-facts you can go and collect. "Is this vendor capable of delivering the contract?" breaks down into registered entity and status, filing history, litigation, adverse media, ownership, and operating footprint.

Finish with the two things people skip: success criteria (what a complete answer looks like, decided in advance so you are not moving the goalposts to fit what you found) and a stop condition (you stop when the requirements are answered to the needed confidence, when more collection stops changing your assessment, or when you hit a legal or ethical wall).

The rule: if you cannot say in one sentence what decision your work will inform and what would change that decision, do not start collecting. Go back to the requester.

2. Collect on purpose, not on reflex

The dominant failure mode in OSINT is "collect everything." It feels productive and it is actively harmful. It buries the signal, burns the one resource you cannot recover (attention), and makes your work impossible to reproduce.

Build a plan that maps each information requirement to source layers, a method, and a cadence. Think in layers, not favorite sites:

Assign a cadence to anything that changes, and treat reproducibility as a requirement, not a nicety. Another competent investigator, handed your plan and your log, should retrace your steps and reach the same material. That means logging queries, dates, and sources as you go, not reconstructing them from memory later.

Collect without burning the operation. Before you touch a source, decide whether the subject can see that you did, and whether that matters. Reading is passive. Following, connecting, or messaging is active, and it can tip the subject and create legal and ethical problems. Some platforms, professional networks especially, tell a user who viewed their profile. Route sensitive research through non-attributable infrastructure. Keep any research account fully separate from your personal identity, created with organizational approval and inside the platform's rules and the law. And pace your collection, because aggressive automated pulls trip anti-bot defenses and can taint or block the whole effort.

Stay inside the law. "Stop when you hit a legal boundary" only works if you know where the boundaries are. The recurring ones: platform terms of service and anti-scraping clauses; computer-misuse exposure (in the US, the Computer Fraud and Abuse Act) for automated collection or pretext access; and, anywhere you operate in or collect on people in the EU or other GDPR-style regimes, a documented lawful basis for processing personal data. Some things are collectible in principle but out of bounds in practice. That call belongs in the plan, confirmed with counsel when the stakes are high, not discovered after the fact.

3. Grade your sources, and never mistake repetition for proof

The most common analytic error in open sources is treating repetition as corroboration. Ten accounts posting the same claim are not ten sources. They may be one source and nine echoes.

Trace every claim back toward its origin. Separate primary sources (the original document, the person who was there, the first upload of an image) from derivative ones (reporting about the document, a re-upload, a screenshot of a screenshot). Claims launder themselves as they travel. A hedged sentence in a forum becomes a confident headline three hops later, with the uncertainty stripped out. Chase it upstream until you hit bedrock or run out of trail, and note where it ended. Ask who benefits from the claim being believed. Motive does not make a claim false, but it tells you how hard to push.

To keep source assessment consistent instead of impressionistic, grade the source and the information on two separate scales. This is the A-F and 1-6 matrix in US Army doctrine (FM 2-22.3, Appendix B), known in British and NATO practice as the Admiralty Code.

Source reliability, how much you trust the source:

Information credibility, how much you trust the specific claim:

A datum tagged "B2" tells a downstream reader far more than "per an online post." But do not oversell the matrix. It is only as good as the discipline behind it. Studies of how analysts use it find grades cluster on the middle values, and the two scales, which are meant to be independent, get treated as if they move together. A "B2" is shorthand for your reasoning, never a replacement for showing the actual sourcing chain.

It also helps to name what you are looking at. Claire Wardle and Hossein Derakhshan's Information Disorder framework separates misinformation (false, no intent to harm), disinformation (false, intent to harm), and malinformation (true information weaponized to harm). The one that catches experienced people is "false context," genuine media presented as something it is not, because the media itself is real.

4. Reasoning is the actual skill

Collection gives you clues. Clues are not claims, and claims are not conclusions. The move between them is reasoning, and that is where most failure happens.

Watch for the "therefore trap," a chain of individually plausible steps that quietly sheds probability at every link and lands on a confident conclusion built from multiplied uncertainty. "The account posts in this time zone, therefore the user lives there, therefore they were at the event, therefore they are involved." Every "therefore" is a hinge, and every hinge can be wrong. Make the chain explicit and test each link on its own: is this an observation or an inference, and what else could produce the same clue?

Two structures keep reasoning honest. First, hypothesis trees: list the plausible explanations, including the boring ones (coincidence, error, staging, unrelated activity), before you commit. If you have only one hypothesis, you are not analyzing, you are confirming. Second, evidence-hypothesis matrices: lay your evidence against each hypothesis and look for what is inconsistent, because a single solid inconsistency can push a candidate to the bottom of the list, while consistent evidence rarely proves anything. That is the core of Richards Heuer's Analysis of Competing Hypotheses. One caution the method itself demands: before you drop a hypothesis on one inconsistency, ask whether that inconsistent fact could itself be wrong, mislabeled, or planted. Any single piece of evidence can be the deception.

5. Use a few structured techniques as fast habits

Structured analytic techniques sound bureaucratic. A handful of them are just fast, high-leverage habits that slow your intuitive fast thinking long enough for your deliberate thinking to check it. Five earn their keep:

None of these takes long. A Premortem is one question said out loud: "It is six months from now and we were badly wrong. What did we miss?" The discipline is doing them at all, and doing them before you are committed to an answer instead of after you have to defend one.

6. Geolocate and chronolocate by layered, falsifiable evidence

Placing an image or video in space and time is where OSINT reasoning becomes visible and testable. Do it as an accumulation of independent evidence, not a single "gotcha" match.

Work the layers, each one narrowing the possibilities: terrain and skyline; the built environment (architecture, road markings, signage, utility poles, guardrails); language and text (scripts, dialect, phone and plate formats); vegetation and climate; region-specific vehicles and objects; and sun and shadow. Then match those constraints against satellite and street-level imagery, the workflow Bellingcat has documented and maintained for a decade.

Two disciplines separate credible geolocation from confirmation bias:

Try to prove it wrong. If a video is said to be in City A, hunt for what should be in City A and is absent, or present and inconsistent. A location you failed to disprove is far stronger than one you merely matched.

Reason carefully about time. Shadow direction and length, with a solar calculator like SunCalc, can constrain the time of day and corroborate or contradict a claimed timestamp. Be precise: shadow-length math needs a reference object of known height; calculators return local solar time, so convert for time zone and daylight saving before comparing to a clock timestamp; and use the shadow's direction to resolve morning versus afternoon.

Treat metadata with caution, not faith. EXIF timestamps and GPS tags are routinely stripped on upload, are trivially editable, and carry time-zone ambiguity worth hours of error. Metadata is a lead, not proof. The strongest work rests on independent corroboration: two or more captures, from different uploaders and angles, that agree on the same scene. The Berkeley Protocol, the leading international standard for this, is explicit that verification, not metadata, carries the weight.

7. Assume some of it was planted, and watch your own head

The modern information environment is adversarial by default. It includes staged media, recycled footage, imposter accounts, and coordinated inauthentic behavior built to manufacture the appearance of consensus.

Indicators worth checking, none conclusive alone: lighting or shadows that do not agree, edges and artifacts around inserted elements, audio that does not match the scene; reverse image results that show the same media predating the claimed event; and clusters of accounts posting near-identical content in tight windows, with bunched creation dates and thin histories. When you reverse-search an image, use more than one engine, because none is comprehensive. Google, Yandex, and TinEye behave very differently. Yandex is often strongest on faces and places, and TinEye is best at finding the earliest known copy.

Synthetic media deserves its own checklist. Generative models now produce photoreal images, cloned voices, and video that the classic manipulation tells above will miss, because nothing was spliced. The whole frame was synthesized. Look for the model's own artifacts: implausible hands, teeth, and jewelry; garbled text on signs and clothing; reflections and shadows that do not agree; backgrounds that warp; and, in video, flicker, morphing edges, and unnatural blinking or lip-sync. Treat the emerging C2PA Content Credentials provenance standard as a positive signal when present, and understand that its absence proves nothing. And do not lean on automated detectors as proof. They are unreliable in both directions, false-positive on real media and false-negative on synthetic. Apply the same falsification discipline in reverse: do not claim "this is AI-generated" without corroboration any more than you would claim a location without it. Over-claiming is as damaging on cross as missing it.

The hardest adversary is internal. The biggest threats to your analysis are your own habits: confirmation bias (favoring what fits your theory), anchoring (the first framing dragging every later estimate), availability (treating the vivid and recent as more probable), groupthink, and overconfidence. You do not beat these by resolving to be objective. You instrument against them: write assumptions down before you look, work by disconfirmation, assign someone to argue the other side, and run a Premortem. Structure is the countermeasure. Willpower is not.

8. Write for the reader with thirty seconds and the adversary with all day

A finding a decision-maker cannot act on, or that collapses under cross-examination, was not worth collecting.

The section most guides skip: capture and chain of custody

Online evidence is perishable and mutable. A post gets edited or deleted, a page changes, an account vanishes, all between the moment you see it and the moment anyone asks you to prove it. If there is any chance your finding reaches a report, a regulator, or a courtroom, you preserve at the moment of collection, not later. This is the section that decides whether your work is usable at all.

Preserve so the artifact is authentic (it is what you say it is) and its provenance is traceable (where it came from and when you got it). A minimum routine:

For anything court-bound, automated contemporaneous capture is the standard, not an ad hoc routine. Purpose-built tools (Hunchly is the common one) silently log, hash, and timestamp every page you visit and export a case report, which is far harder to attack than a folder of screenshots assembled afterward. For high-stakes work, screen-record the session and note the tool version.

Know the limits of what this buys you. Authentication and integrity are the two preconditions you control by capturing well. They are necessary, not sufficient. Admissibility also turns on rules you do not control and that vary by jurisdiction: authentication standards (in US federal practice, Federal Rules of Evidence 901 and 902), the hearsay problem when a post is offered for the truth of what it says, relevance, and best-evidence rules. In the EU and other GDPR-influenced systems, data-protection and civil-evidence rules differ, and evidence gathered without a lawful basis can be excluded. You are not making those calls, and none of this is legal advice. You are making sure that when counsel does make them, your record supports the case instead of sinking it. The discipline costs minutes at collection and is impossible to reconstruct later. Do it every time, or you have a story instead of evidence.

A field checklist worth saving

Run it end to end on anything that might be seen by someone other than you.

Before collection:

During collection:

During analysis:

Before delivery:

Bottom line

The tools will keep changing. The judgment is the durable asset, and it is what decides whether the work holds up when it matters. Frame the question before you touch a tool. Plan collection, and collect without burning the operation or breaking the law. Never mistake repetition for corroboration. Reason by disconfirmation. Geolocate by falsifiable evidence and treat metadata as a lead. Assume some of it was planted, and treat synthetic media as a first-order threat. Write the bottom line first. And capture, hash, and archive at the moment of collection, because evidence you did not preserve is just a memory.

Take it further

Everything above is the what. Turning it into a repeatable method that holds up in a report, a declaration, a regulatory filing, or a courtroom is the how, and that is the entire focus of the Certified OSINT Investigator, Court-Ready Practitioner (COI-CRP) course from The Waldrep Company.

It is self-paced and online: 10 modules, 80 lessons, 8 downloadable templates, 10 module quizzes, and a 100-question final exam (75% to pass, unlimited free retakes). It covers the full arc of this newsletter and more: mission framing and investigation discipline, the legal frame for OSINT (Fourth Amendment, CFAA and the public-interface rule, GDPR, cross-border collection and MLAT, lawful use of breach data), operational security and identity resolution, social media intelligence, image and geolocation analysis, dark web and cryptocurrency tracing, AI-assisted OSINT with a hallucination-control discipline, chain of custody and hashing for every artifact, court-ready reporting, and Daubert, Frye, and FRE 702. It ends with a mock cross-examination on every technique in the course.

Tuition is $2,497. Government purchase orders and net-30 invoicing are accepted, and agency licensing is available at 10, 25, and unlimited seats.

Enroll or download the full syllabus, a free PDF, at thewaldrepcompany.com/courses/coi-crp/. For agency or purchase-order enrollment, call (251) 216-1164.

If this was useful, subscribe to Digital Forensics Today below and pass it to an investigator, analyst, or attorney who works with open-source evidence. Every issue is about doing OSINT and digital forensics that stands up when someone asks how you know.

OSINT Digital Forensics Investigations Open Source Intelligence Intelligence Analysis

Sources and further reading