Every open-source investigation eventually meets the same question, and it is the one that decides whether the work was worth anything: how do you know?
Anyone with a browser can find something. The hard part, the part that separates an investigator from a search-engine operator, is turning what you found into a claim you can defend when a hostile attorney, a skeptical editor, or a decision-maker with real consequences on the line pushes back.
Investigations tend to fail on that question in the same few ways. Someone matches an image to a location instead of trying to disprove it, and matches the wrong one. Someone finds a decisive post and screenshots it a week later, after it has been edited, with nothing preserved. The finding was real. The proof was not.
OSINT gets taught as a tour of tools and websites. That is backwards. Tools change every quarter. Platforms lock down, sites die, and this year's clever trick is next year's dead link. What lasts is judgment: framing a question, planning collection, weighing sources, reasoning under uncertainty, guarding against bias, and writing findings that survive scrutiny.
What follows is that discipline in eight practices, plus the one section most field guides skip. Every example is generic. The methods are not.
1. Frame the mission before you open a browser
Most bad investigations are lost before collection starts, because someone accepted a vague tasking and never turned it into an answerable question. "Find everything on this company" is not a question. No boundary, no success criterion, no stop condition. It produces a pile of material and no judgment.
Sort what you have into three buckets first:
- Known: facts you can already assert with a source.
- Assumed: things you are treating as true without proof, usually without noticing. This is where investigations quietly go wrong. Write them down.
- Unknown: the specific gaps that, if filled, would change your answer.
Then define your Priority Intelligence Requirements, the two or three decision-driving questions your customer actually needs answered, and under each one the Essential Elements of Information, the concrete sub-facts you can go and collect. "Is this vendor capable of delivering the contract?" breaks down into registered entity and status, filing history, litigation, adverse media, ownership, and operating footprint.
Finish with the two things people skip: success criteria (what a complete answer looks like, decided in advance so you are not moving the goalposts to fit what you found) and a stop condition (you stop when the requirements are answered to the needed confidence, when more collection stops changing your assessment, or when you hit a legal or ethical wall).
The rule: if you cannot say in one sentence what decision your work will inform and what would change that decision, do not start collecting. Go back to the requester.
2. Collect on purpose, not on reflex
The dominant failure mode in OSINT is "collect everything." It feels productive and it is actively harmful. It buries the signal, burns the one resource you cannot recover (attention), and makes your work impossible to reproduce.
Build a plan that maps each information requirement to source layers, a method, and a cadence. Think in layers, not favorite sites:
- Surface: search engines, mainstream and local media, official statements.
- Registry and record: corporate registries, court dockets, property and regulatory filings, sanctions lists.
- Platform and social: profiles, posts, connections, and the metadata around them.
- Technical: domains, WHOIS and DNS history, certificate transparency, infrastructure links.
- Media: images, video, and their provenance.
- Archival: cached and archived versions of all of the above.
Assign a cadence to anything that changes, and treat reproducibility as a requirement, not a nicety. Another competent investigator, handed your plan and your log, should retrace your steps and reach the same material. That means logging queries, dates, and sources as you go, not reconstructing them from memory later.
Collect without burning the operation. Before you touch a source, decide whether the subject can see that you did, and whether that matters. Reading is passive. Following, connecting, or messaging is active, and it can tip the subject and create legal and ethical problems. Some platforms, professional networks especially, tell a user who viewed their profile. Route sensitive research through non-attributable infrastructure. Keep any research account fully separate from your personal identity, created with organizational approval and inside the platform's rules and the law. And pace your collection, because aggressive automated pulls trip anti-bot defenses and can taint or block the whole effort.
Stay inside the law. "Stop when you hit a legal boundary" only works if you know where the boundaries are. The recurring ones: platform terms of service and anti-scraping clauses; computer-misuse exposure (in the US, the Computer Fraud and Abuse Act) for automated collection or pretext access; and, anywhere you operate in or collect on people in the EU or other GDPR-style regimes, a documented lawful basis for processing personal data. Some things are collectible in principle but out of bounds in practice. That call belongs in the plan, confirmed with counsel when the stakes are high, not discovered after the fact.
3. Grade your sources, and never mistake repetition for proof
The most common analytic error in open sources is treating repetition as corroboration. Ten accounts posting the same claim are not ten sources. They may be one source and nine echoes.
Trace every claim back toward its origin. Separate primary sources (the original document, the person who was there, the first upload of an image) from derivative ones (reporting about the document, a re-upload, a screenshot of a screenshot). Claims launder themselves as they travel. A hedged sentence in a forum becomes a confident headline three hops later, with the uncertainty stripped out. Chase it upstream until you hit bedrock or run out of trail, and note where it ended. Ask who benefits from the claim being believed. Motive does not make a claim false, but it tells you how hard to push.
To keep source assessment consistent instead of impressionistic, grade the source and the information on two separate scales. This is the A-F and 1-6 matrix in US Army doctrine (FM 2-22.3, Appendix B), known in British and NATO practice as the Admiralty Code.
Source reliability, how much you trust the source:
- A: Reliable
- B: Usually reliable
- C: Fairly reliable
- D: Not usually reliable
- E: Unreliable
- F: Cannot be judged
Information credibility, how much you trust the specific claim:
- 1: Confirmed by independent sources
- 2: Probably true
- 3: Possibly true
- 4: Doubtfully true
- 5: Improbable
- 6: Cannot be judged
A datum tagged "B2" tells a downstream reader far more than "per an online post." But do not oversell the matrix. It is only as good as the discipline behind it. Studies of how analysts use it find grades cluster on the middle values, and the two scales, which are meant to be independent, get treated as if they move together. A "B2" is shorthand for your reasoning, never a replacement for showing the actual sourcing chain.
It also helps to name what you are looking at. Claire Wardle and Hossein Derakhshan's Information Disorder framework separates misinformation (false, no intent to harm), disinformation (false, intent to harm), and malinformation (true information weaponized to harm). The one that catches experienced people is "false context," genuine media presented as something it is not, because the media itself is real.
4. Reasoning is the actual skill
Collection gives you clues. Clues are not claims, and claims are not conclusions. The move between them is reasoning, and that is where most failure happens.
Watch for the "therefore trap," a chain of individually plausible steps that quietly sheds probability at every link and lands on a confident conclusion built from multiplied uncertainty. "The account posts in this time zone, therefore the user lives there, therefore they were at the event, therefore they are involved." Every "therefore" is a hinge, and every hinge can be wrong. Make the chain explicit and test each link on its own: is this an observation or an inference, and what else could produce the same clue?
Two structures keep reasoning honest. First, hypothesis trees: list the plausible explanations, including the boring ones (coincidence, error, staging, unrelated activity), before you commit. If you have only one hypothesis, you are not analyzing, you are confirming. Second, evidence-hypothesis matrices: lay your evidence against each hypothesis and look for what is inconsistent, because a single solid inconsistency can push a candidate to the bottom of the list, while consistent evidence rarely proves anything. That is the core of Richards Heuer's Analysis of Competing Hypotheses. One caution the method itself demands: before you drop a hypothesis on one inconsistency, ask whether that inconsistent fact could itself be wrong, mislabeled, or planted. Any single piece of evidence can be the deception.
5. Use a few structured techniques as fast habits
Structured analytic techniques sound bureaucratic. A handful of them are just fast, high-leverage habits that slow your intuitive fast thinking long enough for your deliberate thinking to check it. Five earn their keep:
- Key Assumptions Check. Surfaces the load-bearing beliefs you are treating as fact. Run it at the start, and any time the answer feels obvious.
- Analysis of Competing Hypotheses. Scores evidence against rival explanations to find disconfirmers. Use it when more than one story fits and the stakes are high.
- Quality of Information Check. Re-audits your sourcing for reliability, currency, and laundering. Do it before you write, and before any number goes into a report.
- Red Team or Devil's Advocacy. Argues the opposing case in good faith. Reach for it when consensus formed fast or nobody in the room disagrees.
- Premortem. Assume the assessment was wrong, then explain how. Ask it just before delivery to catch overconfidence.
None of these takes long. A Premortem is one question said out loud: "It is six months from now and we were badly wrong. What did we miss?" The discipline is doing them at all, and doing them before you are committed to an answer instead of after you have to defend one.
6. Geolocate and chronolocate by layered, falsifiable evidence
Placing an image or video in space and time is where OSINT reasoning becomes visible and testable. Do it as an accumulation of independent evidence, not a single "gotcha" match.
Work the layers, each one narrowing the possibilities: terrain and skyline; the built environment (architecture, road markings, signage, utility poles, guardrails); language and text (scripts, dialect, phone and plate formats); vegetation and climate; region-specific vehicles and objects; and sun and shadow. Then match those constraints against satellite and street-level imagery, the workflow Bellingcat has documented and maintained for a decade.
Two disciplines separate credible geolocation from confirmation bias:
Try to prove it wrong. If a video is said to be in City A, hunt for what should be in City A and is absent, or present and inconsistent. A location you failed to disprove is far stronger than one you merely matched.
Reason carefully about time. Shadow direction and length, with a solar calculator like SunCalc, can constrain the time of day and corroborate or contradict a claimed timestamp. Be precise: shadow-length math needs a reference object of known height; calculators return local solar time, so convert for time zone and daylight saving before comparing to a clock timestamp; and use the shadow's direction to resolve morning versus afternoon.
Treat metadata with caution, not faith. EXIF timestamps and GPS tags are routinely stripped on upload, are trivially editable, and carry time-zone ambiguity worth hours of error. Metadata is a lead, not proof. The strongest work rests on independent corroboration: two or more captures, from different uploaders and angles, that agree on the same scene. The Berkeley Protocol, the leading international standard for this, is explicit that verification, not metadata, carries the weight.
7. Assume some of it was planted, and watch your own head
The modern information environment is adversarial by default. It includes staged media, recycled footage, imposter accounts, and coordinated inauthentic behavior built to manufacture the appearance of consensus.
Indicators worth checking, none conclusive alone: lighting or shadows that do not agree, edges and artifacts around inserted elements, audio that does not match the scene; reverse image results that show the same media predating the claimed event; and clusters of accounts posting near-identical content in tight windows, with bunched creation dates and thin histories. When you reverse-search an image, use more than one engine, because none is comprehensive. Google, Yandex, and TinEye behave very differently. Yandex is often strongest on faces and places, and TinEye is best at finding the earliest known copy.
Synthetic media deserves its own checklist. Generative models now produce photoreal images, cloned voices, and video that the classic manipulation tells above will miss, because nothing was spliced. The whole frame was synthesized. Look for the model's own artifacts: implausible hands, teeth, and jewelry; garbled text on signs and clothing; reflections and shadows that do not agree; backgrounds that warp; and, in video, flicker, morphing edges, and unnatural blinking or lip-sync. Treat the emerging C2PA Content Credentials provenance standard as a positive signal when present, and understand that its absence proves nothing. And do not lean on automated detectors as proof. They are unreliable in both directions, false-positive on real media and false-negative on synthetic. Apply the same falsification discipline in reverse: do not claim "this is AI-generated" without corroboration any more than you would claim a location without it. Over-claiming is as damaging on cross as missing it.
The hardest adversary is internal. The biggest threats to your analysis are your own habits: confirmation bias (favoring what fits your theory), anchoring (the first framing dragging every later estimate), availability (treating the vivid and recent as more probable), groupthink, and overconfidence. You do not beat these by resolving to be objective. You instrument against them: write assumptions down before you look, work by disconfirmation, assign someone to argue the other side, and run a Premortem. Structure is the countermeasure. Willpower is not.
8. Write for the reader with thirty seconds and the adversary with all day
A finding a decision-maker cannot act on, or that collapses under cross-examination, was not worth collecting.
- Lead with the bottom line. State the answer first, then support it. This is the Army writing standard: main point up front, active voice, understood in a single reading. Do not make a busy reader dig your conclusion out of paragraph nine.
- Build every assertion as claim, evidence, confidence. The claim, the specific source behind it, and a calibrated confidence level. A claim with no evidence is an opinion. Evidence with no confidence level makes the reader guess how far to trust it.
- Calibrate with standardized language, and separate two things readers constantly blur: how likely something is, and how confident you are in that judgment. The US intelligence community's ICD 203 gives a fixed likelihood ladder (from almost no chance up to almost certain) and a separate confidence level (low, moderate, high) driven by the quality of your sourcing. Sherman Kent made this case in 1964: the word "probable" hides wildly different odds in different readers' heads. Pick a scale and use it consistently.
- Keep the language neutral, and name your gaps. Strip loaded adjectives and rhetorical certainty. Say plainly what you could not verify. Naming a gap builds credibility. Hiding it destroys credibility the moment it surfaces.
The section most guides skip: capture and chain of custody
Online evidence is perishable and mutable. A post gets edited or deleted, a page changes, an account vanishes, all between the moment you see it and the moment anyone asks you to prove it. If there is any chance your finding reaches a report, a regulator, or a courtroom, you preserve at the moment of collection, not later. This is the section that decides whether your work is usable at all.
Preserve so the artifact is authentic (it is what you say it is) and its provenance is traceable (where it came from and when you got it). A minimum routine:
- Screenshot with context. Capture the full page including the URL bar, and a full-page or scrolling capture, not a cropped fragment. A visible clock in the corner is a convenience, not proof, because it is trivially spoofed. The weight comes from the archive and the log below.
- Preserve the source, not just the picture. Save the underlying HTML, and for video the original file where lawful, so you hold more than a rendered image.
- Hash every capture. Compute a cryptographic hash (for example SHA-256) at collection time and record it. A matching hash later shows the file has not changed since.
- Archive independently. Push the live URL to a third-party archive such as the Internet Archive or archive.today, so a timestamped copy exists outside your machine.
- Log the collection itself. For every artifact, record the URL, the date and time with time zone, the tool or method, and the collector. That log is your chain of custody.
For anything court-bound, automated contemporaneous capture is the standard, not an ad hoc routine. Purpose-built tools (Hunchly is the common one) silently log, hash, and timestamp every page you visit and export a case report, which is far harder to attack than a folder of screenshots assembled afterward. For high-stakes work, screen-record the session and note the tool version.
Know the limits of what this buys you. Authentication and integrity are the two preconditions you control by capturing well. They are necessary, not sufficient. Admissibility also turns on rules you do not control and that vary by jurisdiction: authentication standards (in US federal practice, Federal Rules of Evidence 901 and 902), the hearsay problem when a post is offered for the truth of what it says, relevance, and best-evidence rules. In the EU and other GDPR-influenced systems, data-protection and civil-evidence rules differ, and evidence gathered without a lawful basis can be excluded. You are not making those calls, and none of this is legal advice. You are making sure that when counsel does make them, your record supports the case instead of sinking it. The discipline costs minutes at collection and is impossible to reconstruct later. Do it every time, or you have a story instead of evidence.
A field checklist worth saving
Run it end to end on anything that might be seen by someone other than you.
Before collection:
- State the decision the work informs and what would change it
- Write the intelligence requirements; separate known, assumed, and unknown
- Set success criteria and a stop condition
- Confirm the legal and ethical boundaries, and how to collect without tipping the subject
During collection:
- Work a plan that maps requirements to source layers, method, and cadence
- Collect passively and non-attributably unless a case decision says otherwise
- Log every query, source, URL, and capture time as you go
- Capture, hash, and archive at the moment of collection
- Trace each key claim to a primary source, and note where the trail ends
During analysis:
- Enumerate competing hypotheses, not just the favored one
- Run a Key Assumptions Check, and ACH where the stakes warrant
- Grade sources and information, and keep likelihood separate from confidence
- Check media for both classic manipulation and synthetic generation
- Run a Red Team pass or Premortem before committing
Before delivery:
- State the bottom line; make every claim carry evidence and calibrated confidence
- Keep language neutral; name gaps explicitly
- Annex evidence with URLs, timestamps, hashes, and method
- Confirm a second investigator could reproduce the work from the log
Bottom line
The tools will keep changing. The judgment is the durable asset, and it is what decides whether the work holds up when it matters. Frame the question before you touch a tool. Plan collection, and collect without burning the operation or breaking the law. Never mistake repetition for corroboration. Reason by disconfirmation. Geolocate by falsifiable evidence and treat metadata as a lead. Assume some of it was planted, and treat synthetic media as a first-order threat. Write the bottom line first. And capture, hash, and archive at the moment of collection, because evidence you did not preserve is just a memory.
Take it further
Everything above is the what. Turning it into a repeatable method that holds up in a report, a declaration, a regulatory filing, or a courtroom is the how, and that is the entire focus of the Certified OSINT Investigator, Court-Ready Practitioner (COI-CRP) course from The Waldrep Company.
It is self-paced and online: 10 modules, 80 lessons, 8 downloadable templates, 10 module quizzes, and a 100-question final exam (75% to pass, unlimited free retakes). It covers the full arc of this newsletter and more: mission framing and investigation discipline, the legal frame for OSINT (Fourth Amendment, CFAA and the public-interface rule, GDPR, cross-border collection and MLAT, lawful use of breach data), operational security and identity resolution, social media intelligence, image and geolocation analysis, dark web and cryptocurrency tracing, AI-assisted OSINT with a hallucination-control discipline, chain of custody and hashing for every artifact, court-ready reporting, and Daubert, Frye, and FRE 702. It ends with a mock cross-examination on every technique in the course.
Tuition is $2,497. Government purchase orders and net-30 invoicing are accepted, and agency licensing is available at 10, 25, and unlimited seats.
Enroll or download the full syllabus, a free PDF, at thewaldrepcompany.com/courses/coi-crp/. For agency or purchase-order enrollment, call (251) 216-1164.
If this was useful, subscribe to Digital Forensics Today below and pass it to an investigator, analyst, or attorney who works with open-source evidence. Every issue is about doing OSINT and digital forensics that stands up when someone asks how you know.
Sources and further reading
- Richards J. Heuer, Jr., Psychology of Intelligence Analysis, CIA Center for the Study of Intelligence, 1999. cia.gov
- A Tradecraft Primer: Structured Analytic Techniques for Improving Intelligence Analysis, CIA, 2009. cia.gov (PDF)
- Randolph H. Pherson and Richards J. Heuer, Jr., Structured Analytic Techniques for Intelligence Analysis, 3rd ed., CQ Press, 2019. us.sagepub.com
- Office of the Director of National Intelligence, Intelligence Community Directive (ICD) 203: Analytic Standards, 2023. dni.gov (PDF)
- Sherman Kent, "Words of Estimative Probability," Studies in Intelligence, 1964. cia.gov
- UN OHCHR and UC Berkeley Human Rights Center, Berkeley Protocol on Digital Open Source Investigations, United Nations, 2022. ohchr.org
- Bellingcat, Online Investigation Toolkit, 2024. bellingcat.gitbook.io/toolkit
- Eliot Higgins, "A Beginner's Guide to Geolocating Videos," Bellingcat, 2014. bellingcat.com
- Michael Bazzell and Jason Edison, OSINT Techniques, 11th ed., 2024. inteltechniques.com
- Claire Wardle and Hossein Derakhshan, Information Disorder, Council of Europe, 2017. rm.coe.int
- Daniel Kahneman, Thinking, Fast and Slow, Farrar, Straus and Giroux, 2011. ISBN 9780374275631.
- Headquarters, Department of the Army, FM 2-22.3, Human Intelligence Collector Operations, Appendix B, 2006. irp.fas.org (PDF)
- Headquarters, Department of the Army, Army Regulation 25-50, Preparing and Managing Correspondence, 2020. armypubs.army.mil (PDF)