OSINT
OSINT, Attribution, and the Honest Expression of Confidence
Open sources can take you a long way toward naming an adversary — and they can take you confidently to the wrong name. The discipline is not in the collection; it is in what you are willing to claim from it.
- OSINT
- attribution
- confidence
- analytic standards
- bias
Attribution is the most consequential judgement in cyber threat intelligence and the one most often made badly. It is consequential because organisations, and occasionally governments, act on it. It is made badly because the evidence available from open sources is abundant, cheap, and systematically biased — a combination that produces confident conclusions far faster than it produces correct ones.
This article is about the gap between those two, and about the specific practices that narrow it.
Three questions that get called “attribution”
The word covers three distinct problems of increasing difficulty, and conflating them is the first error.
Clustering — are these events the same activity? The lowest bar and the most tractable. Given two intrusions, do they share infrastructure, tooling, timing, targeting, or implementation detail sufficient to treat them as one campaign? This is a similarity judgement and it can be made defensibly from technical evidence alone.
Naming — is this activity the group already tracked as X? Harder, because it inherits every assumption in the existing cluster. When you say “this is APT28” you are asserting that your new evidence belongs in a bucket that someone else assembled, under criteria you may not have inspected, from evidence you may not have seen. Vendor group names are not interchangeable; two vendors’ clusters bearing “the same” actor’s name routinely differ in membership.
Responsibility — which organisation or state directed this? Hardest by a wide margin, and generally not answerable from open sources. It requires establishing a chain from operator to sponsor, and open evidence almost never spans that chain. Most public “attribution” that reaches this level is either drawing on non-public collection it cannot cite, or overreaching.
A great deal of analytic confusion dissolves the moment a report states which of these three it is claiming. Reports frequently present clustering evidence and draw a responsibility conclusion, with the transition happening silently in a single sentence.
Why open sources bias toward the wrong answer
OSINT’s defining advantage — everything is available — is also its defining hazard. Four biases are structural rather than accidental, which means discipline, not effort, is the remedy.
Survivorship in what gets published
You see the adversaries who were caught, written about, and had infrastructure documented. The ones who were never noticed have no footprint in your evidence base at all. Any frequency claim derived from published reporting — “this technique is used mainly by X” — is a claim about reporting, not about activity, and the two diverge in exactly the way that matters.
Circular reporting
An open-source claim gets republished, aggregated, translated, cited, and eventually appears in five apparently independent places. Corroboration requires independent origin, not independent publication. Without provenance tracked to the original source, an analyst reliably mistakes one claim repeated five times for five claims agreeing.
This is the single most common way OSINT confidence inflates without any new evidence existing. The correct discipline is mechanical: record, for every claim, the earliest source you can reach, and treat downstream copies as zero additional weight.
Deliberate deception
The adversary can read your sources. Infrastructure can be registered to invite a particular inference; artefacts can carry language strings, timezone stamps, compile timestamps, and code comments chosen to point elsewhere. These have all been observed in real operations. The asymmetry is severe: planting misleading evidence is cheap, and detecting that evidence was planted is expensive.
The defensive posture is not paranoia but differential weighting — evidence that is cheap for the adversary to forge should carry less weight than evidence that is costly. A Cyrillic resource string is nearly free. Consistent operational tempo across two years of activity, aligned to a working week and holiday calendar, is not.
Confirmation and anchoring
The first hypothesis formed shapes every subsequent search. In OSINT the effect is amplified because search is query-driven: you find what you look for, and what you look for is what you already suspect. An analyst who begins with “this looks like X” will assemble a convincing case for X from a corpus that also contains a convincing case for Y.
The standard countermeasure is Analysis of Competing Hypotheses: enumerate the plausible hypotheses first, then score evidence by its diagnosticity — how well it discriminates between hypotheses, not how consistent it is with the favoured one. Evidence consistent with every hypothesis has no diagnostic value regardless of how compelling it feels. This inverts the natural workflow, which is precisely the point.
The evidence hierarchy, ordered by cost to the adversary
Attribution evidence is not uniform, and its weight should follow how expensive it is to falsify.
Weak — cheap to forge or coincidental Language strings in binaries · compile timestamps · timezone settings · code comments · registrant names in WHOIS · a single shared IP on bulletproof or shared hosting · reuse of publicly available tooling.
Moderate — requires effort but achievable Distinctive implementation quirks in custom code · certificate reuse across campaigns · operational tempo consistent with a specific working calendar · targeting that is coherent with a specific set of interests · shared C2 protocol idiosyncrasies.
Strong — expensive, and hard to fake consistently over time Unique custom tooling with non-trivial shared code lineage · operational security failures revealing consistent operator behaviour · infrastructure management patterns sustained across years · correlation of targeting with a specific decision-making calendar · convergence of several independent lines that do not share a collection path.
That last phrase is the one that carries the weight, and it is the most frequently violated. Two findings derived from the same source are one finding. If your infrastructure analysis and your malware analysis both trace back to the same sandbox report, they are not corroborating lines of evidence; they are one line of evidence displayed twice. Independence must be established at the level of origin, not at the level of analysis.
Stating confidence so that it survives contact with the reader
An assessment with no stated confidence is not modest; it is unfalsifiable. The reader will assign a confidence anyway, and they will get it wrong in whichever direction the prose leans.
Two public frameworks do most of the work.
Admiralty Code grades the source and the content on separate axes — reliability A–F, credibility 1–6. The separation matters because a highly reliable source can report something implausible, and an unreliable source can be right. A commercial “confidence score” that collapses both into one number has destroyed the distinction that tells an analyst whether to go looking for corroboration.
ICD 203 fixes the meaning of estimative language so that likely means the same to writer and reader. What matters is not the specific bands but that they are published and fixed. Unstandardised, “possible” and “probable” are read at wildly different probabilities by different readers — a miscommunication the producing organisation has no way to detect.
Three states, not two
Most tooling supports confirmed and not confirmed, and that is one state short. There are three:
- Assessed true, with stated confidence — we looked and formed a judgement.
- Assessed false, with stated confidence — we looked and the evidence is against it.
- Unable to assess — we did not look, could not look, or the evidence does not bear on the question.
The third is not “low confidence”. Low confidence is a judgement. Unable to assess is the absence of one. A system with no way to express it will render it as “nothing found”, and “we did not look” and “we looked and found nothing” are opposite claims. Any pipeline that maps missing evidence onto a negative finding is generating false negatives at a rate it cannot measure — because the very records that would reveal the rate are the ones being silently converted.
Practices that actually change the error rate
These are unglamorous and they work.
Record provenance at collection, not at write-up. Source URL, retrieval timestamp, archive copy, and the earliest origin you could reach. Provenance reconstructed later is provenance invented later. Open sources are mutable and disappear; an assessment that cannot be re-examined in a year is an assertion, not a finding.
Enumerate hypotheses before assembling evidence. Including at least one hypothesis you consider unlikely, and including “this is deliberate misdirection” as a standing candidate.
Score evidence for diagnosticity, not consistency. The useful question is never “does this fit?” It is “which hypotheses does this rule out?”
Separate the clustering claim from the naming claim from the responsibility claim. State which one you are making, in the assessment sentence itself.
Write the falsifier. Every assessment should name what observation would overturn it. This is the discipline that most reliably separates analysis from advocacy — an assessment that cannot name its own refutation was not reasoned to, it was arrived at.
Date the assessment and revisit it. Attribution decays. Infrastructure is reassigned, tooling leaks and gets adopted by unrelated actors, groups split and merge. An assessment made from open sources eighteen months ago is a historical artefact, and continuing to cite it as current is a failure of hygiene rather than of analysis.
The uncomfortable conclusion
Good OSINT attribution frequently ends at: “we assess with moderate confidence that this activity clusters with previously reported activity tracked as X; we are unable to assess responsibility from open sources.”
That is a less satisfying sentence than a name and a flag. It is also, very often, the most that the evidence actually supports — and the credibility of an intelligence function is built almost entirely on the occasions when it said so.
สรุปภาษาไทย
การระบุผู้กระทำ (Attribution) คือคำวินิจฉัยที่มีผลกระทบสูงที่สุดในงานข่าวกรอง ภัยคุกคาม และเป็นคำวินิจฉัยที่ทำผิดพลาดบ่อยที่สุด เพราะหลักฐานจากแหล่งเปิดนั้น มีมาก ราคาถูก และมีอคติเชิงโครงสร้าง — ส่วนผสมที่ผลิต ข้อสรุปที่มั่นใจ ได้เร็วกว่า ข้อสรุปที่ถูกต้อง มาก
- สามคำถามที่ถูกเรียกรวมว่า attribution — ① จัดกลุ่ม (เหตุการณ์เหล่านี้เป็นชุดเดียวกันไหม) ② ตั้งชื่อ (ใช่กลุ่ม X ที่ติดตามอยู่ไหม) ③ ความรับผิดชอบ (องค์กรหรือรัฐใดสั่งการ) ⛔ ข้อ ③ แทบไม่มีทางตอบได้จากแหล่งเปิด รายงานจำนวนมากแสดงหลักฐานระดับ ① แล้วสรุประดับ ③ โดยข้ามขั้นในประโยคเดียว
- อคติเชิงโครงสร้างสี่ชนิด — ① อคติผู้รอดชีวิต (เห็นเฉพาะรายที่ถูกจับได้) ② การรายงานวนซ้ำ (ข่าวเดียวถูกอ้างต่อห้าที่ ≠ ห้าแหล่งยืนยันตรงกัน) ③ การลวงโดยเจตนา (ฝ่ายตรงข้ามอ่านแหล่งข่าวของเราได้) ④ อคติยืนยันความเชื่อเดิม
- ถ่วงน้ำหนักหลักฐานตามต้นทุนการปลอม — สตริงภาษา/เขตเวลา/timestamp ปลอมได้แทบฟรี ส่วนจังหวะปฏิบัติการที่สม่ำเสมอตลอดสองปี ปลอมยากกว่ามาก
- 🔑 หลักฐานสองชิ้นที่มาจากแหล่งเดียวกัน คือหลักฐานชิ้นเดียว — ความเป็นอิสระ ต้องพิสูจน์ที่ ต้นกำเนิด ไม่ใช่ที่ ขั้นวิเคราะห์
- สามสถานะ ไม่ใช่สองสถานะ — ประเมินว่าจริง · ประเมินว่าเท็จ · ประเมินไม่ได้ ⛔ “ประเมินไม่ได้” ไม่เท่ากับ “ความเชื่อมั่นต่ำ” ระบบที่ไม่มีช่องนี้จะรายงาน เราไม่ได้ดู ออกมาเป็น เราดูแล้วไม่พบ ซึ่งมีความหมายตรงกันข้าม
- เขียนสิ่งที่จะหักล้างข้อสรุปของตัวเองไว้ด้วยเสมอ — ข้อสรุปที่บอกไม่ได้ว่า อะไรจะล้มมันได้ ไม่ได้มาจากการให้เหตุผล แต่มาจากการเลือกไว้ก่อนแล้ว
บทสรุปที่ไม่สบายใจแต่ถูกต้อง: งาน OSINT ที่ทำอย่างมีวินัยมักจบลงที่ “ประเมินด้วยความเชื่อมั่นปานกลางว่ากิจกรรมนี้จัดอยู่ในชุดเดียวกับที่ติดตามในชื่อ X · ไม่สามารถประเมินความรับผิดชอบได้จากแหล่งเปิด” — ประโยคนี้ให้ความพอใจน้อยกว่าการ ระบุชื่อและธงชาติ แต่ความน่าเชื่อถือของหน่วยข่าวกรองสร้างขึ้นจากโอกาสที่หน่วยนั้น ยอมพูดแบบนี้ เป็นหลัก
References. Richards J. Heuer Jr., Psychology of Intelligence Analysis, CIA Center for the Study of Intelligence (1999); US ODNI Intelligence Community Directive 203, Analytic Standards; NATO STANAG 2511 (Admiralty Code); Sergio Caltagirone, Andrew Pendergast, Christopher Betz, The Diamond Model of Intrusion Analysis (2013); MITRE ATT&CK.