Kelsey Maurine Brickl
Where history exposes power and moral failure

The Hidden Reconstruction Problem in Government and Military AI Systems

Military analyst reviewing surveillance data at a computer workstation. Image © Seventyfour / Adobe Stock, used under Standard License by the author.

Long before ChatGPT, Claude, or Gemini became public-facing products, governments were using artificial intelligence to search data, classify images, screen cargo, and compare records. Police forces compared fingerprints across jurisdictions and checked license plates and facial images against expanding national and transnational databases.

None of this use of AI arrived as science fiction. Adoption was uneven, frequently overpromised, and driven by the belief that software could process material faster than existing personnel and systems. Generative artificial intelligence extends that progression into language itself.

An NSA officer can quickly request an English-language summary of hundreds of intercepted communications initially compiled by France’s DGSE. A military lawyer can request an explanation of competing legal authorities. A diplomat can ask for a comparison of the tone of public statements issued over several months. A physician in a military hospital can organize thousands of clinical notes into a coherent chronology before examining a wounded patient. Thousands of such routine requests are changing how bureaucracies process volumes of information that once required large teams and substantial time. Taken together, they mark an extraordinary shift in how bureaucracies process monumental amounts of information.

The United States has integrated artificial intelligence across numerous civilian and defense functions. Israel’s technology sector has become one of the world’s leading centers for cybersecurity, intelligence technologies, and machine learning research. China continues to invest enormous state resources in artificial intelligence spanning industrial production, surveillance, military modernization, and scientific research. Russia has pursued AI applications in defense, electronic warfare, intelligence analysis, and language processing. NATO members increasingly explore common standards for interoperability, logistics, intelligence support, and defense procurement, while the European Union simultaneously attempts to construct one of the world’s most ambitious regulatory frameworks for artificial intelligence. 

These governments disagree about almost everything that matters geopolitically. Their strategic interests collide. Their legal traditions and alliances differ and their political systems differ even more. Yet each is confronting the same practical question: how should conversational artificial intelligence participate in decisions whose consequences extend far beyond the screen on which an answer appears?

Public debate still concentrates heavily on whether governments should use generative AI at all. As with military aviation and government computing, procurement and operational use are overtaking the abstract debate over adoption. Procurement, institutional need, and strategic competition eventually settled those arguments. Governments adopted the technologies because they solved problems that existing tools could not solve efficiently. The same pattern is unfolding today. The relevant policy question is no longer whether states will employ conversational AI. They already do, and they almost certainly will do so more extensively over the coming decade.

That reality should redirect attention toward a more specific problem. Contemporary evaluation of generative AI focuses overwhelmingly on the answer produced at the end of the interaction. Engineers test hallucination rates, benchmark accuracy, security, bias, information protection, and legal compliance. Those efforts are not just important but critical for efficacy and safety, and many represent genuine advances over earlier generations of artificial intelligence. They nevertheless share a common assumption. They treat the visible response as the primary object worthy of inspection.

Every conversational AI system receives an expression written by a human being. That expression may contain shorthand, slang, errors, omitted context, conflicting evidence, uncertain chronology, or distinctions that depend upon specialist knowledge. Regardless of its quality, the system must first determine what it believes the user is asking before it can produce any answer at all. It cannot skip that step. The reconstruction may occur in fractions of a second, but it remains indispensable. The machine must construct an internal representation of the user’s reasoning, intentions, chronology, relationships, constraints, and evidentiary structure before it can continue.

That intermediate act receives remarkably little public attention despite governing everything that follows. Existing review processes concentrate on outputs, legal consequences, procurement standards, and acceptable-use policies. Yet the computational reconstruction between the prompt and the response remains largely invisible, even though every subsequent answer depends upon it. If that reconstruction subtly changes chronology, certainty, attribution, scope, or relationships among pieces of evidence, an answer may appear perfectly fluent while resting upon reasoning that the human user never intended to provide.

In government, law, diplomacy, intelligence, and military affairs, small changes in how information is reconstructed often matter more than obvious factual errors. A system that invents a nonexistent battle can usually be caught. A system that quietly strengthens uncertainty into confidence, compresses chronology, or shifts the relationship between two documented events presents a more difficult problem because its answer may remain factually plausible while altering the reasoning on which later decisions depend. That is the question governments should begin asking before they become satisfied with answers alone.

Governments are hardly unfamiliar with this problem. In fact, entire professions exist because human beings have always recognized that reconstruction sits between observation and judgment. Historians spend years separating archival evidence from the interpretations imposed upon it by later writers. A witness testifies in court, but the judge or jury must decide what that testimony actually establishes. Intelligence officers distinguish between raw reporting, analytical assessment, and finished intelligence because each stage introduces another opportunity for misunderstanding or distortion. Physicians distinguish what a patient reports from the diagnostic narrative eventually recorded in the medical file. Military commanders routinely receive observations gathered from dozens of independent sources before those observations become a single operational briefing.

None of these disciplines assumes that information passes unchanged from one mind to another, which is why their evidentiary and interpretive procedures exist.

That recognition has shaped institutional procedure for generations. Courts preserve chains of custody. Intelligence agencies classify sources according to reliability and confidence. Academic historians cite archives precisely so that later researchers can revisit the underlying evidence rather than trusting an author’s interpretation. Military organizations document operational timelines because sequence often determines legality as much as outcome. These procedures are expensive, bureaucratically cumbersome, and occasionally frustrating to those who must follow them. Governments nevertheless maintain them because experience has repeatedly demonstrated that subtle changes introduced during reconstruction can produce decisions that diverge dramatically from the underlying facts.

Conversational artificial intelligence has entered this institutional landscape without altering the underlying problem. It has merely automated one stage of it. Every time an analyst asks for a summary of two hundred intelligence reports, the model must determine what constitutes a report, which observations belong together, how uncertainty should be expressed, whether apparently contradictory information should remain contradictory, and how chronology ought to be organized before a single sentence appears on the screen. These judgments occur computationally rather than consciously, but they remain judgments nonetheless. They determine the shape of the response long before anyone evaluates whether that response is persuasive.

Intelligence reporting offers a routine example. Imagine an intelligence analyst reviewing reports from multiple human sources concerning a rapidly developing situation. Several reports contain overlapping observations but disagree about confidence levels. One source believes an event probably occurred. Another considers it merely possible. A third explicitly states that available evidence remains insufficient. A conversational system asked to summarize those materials may produce an elegant paragraph that faithfully includes every major fact while quietly compressing those differences into a smoother narrative. Nothing has been fabricated. No hallucination has occurred. Yet the bureaucratic meaning of the underlying record has shifted because uncertainty itself carried analytical value.

Consider a military legal adviser preparing a memorandum regarding the application of international humanitarian law to a developing operational question. The relevant documents may include intelligence assessments prepared on different days, legal precedents issued years apart, diplomatic communications, and factual updates arriving throughout the day. Chronology is not decorative background. It determines which facts were available when particular decisions were made. A summary that rearranges that sequence while preserving every individual fact may inadvertently alter legal analysis without introducing any objectively false statement.

Medical documentation presents another illustration. Physicians rarely make decisions based upon isolated symptoms. They reconstruct patterns across time. Fever preceding respiratory distress may suggest one diagnosis; respiratory distress preceding fever may suggest another. If a conversational system reorganizes clinical notes into a cleaner narrative while subtly changing sequence, the output may remain grammatically flawless and medically plausible while nevertheless leading physicians toward a different interpretation of the patient’s condition. Modern medicine therefore devotes enormous effort to preserving accurate records rather than attractive summaries.

Diplomacy offers similar examples. Governments communicate through language that often contains deliberate ambiguity. Officials distinguish between what another state has confirmed, what it has suggested privately, what outside observers infer, and what remains unknown. Those distinctions are not stylistic habits inherited from cautious diplomats. They frequently determine how governments assess intent, calculate risk, and respond publicly. A summary that smooths those categories into more confident prose may appear clearer to its reader while simultaneously erasing precisely the uncertainty that policymakers needed to preserve.

These examples share an important characteristic. None depends upon spectacular system failure. No fictional documents appear. No imaginary statutes emerge. No nonexistent military units materialize from thin air. Existing evaluations of generative AI understandably devote substantial attention to preventing those obvious errors because they are visible, measurable, and comparatively straightforward to identify. Reconstruction errors belong to a different category altogether. They arise not from invention but from transformation. The facts often survive. Their relationships do not always survive with equal fidelity.

That contrast deserves considerably more attention than it presently receives because governments increasingly rely upon conversational systems precisely where relationships among facts matter most. Intelligence does not consist merely of information. Law does not consist merely of statutes. Diplomacy does not consist merely of public statements. Military planning does not consist merely of maps and logistics. Every one of those activities depends upon preserving how evidence connects across time, confidence, attribution, and circumstance before officials exercise judgment. The computational reconstruction performed inside a conversational AI therefore becomes more than a technical curiosity. It becomes part of the evidentiary architecture upon which governments increasingly rely.

Governments already require extraordinary levels of verification before acting upon consequential information. Financial auditors reconstruct transactions before certifying accounts. Aviation investigators reconstruct cockpit events before issuing safety recommendations. Intelligence agencies reconstruct events from fragmentary reporting before producing finished assessments. Criminal investigators reconstruct timelines before presenting evidence in court. The common feature is not the profession but the assumption that reconstruction itself deserves scrutiny because later decisions inherit its strengths and weaknesses.

Generative artificial intelligence has introduced another reconstruction layer into that institutional landscape. Unlike the human analyst whose intermediate reasoning may remain visible through notes, drafts, annotations, or discussion, the reconstruction performed inside a conversational model typically disappears before the user ever sees a response. Governments therefore inspect the answer without inspecting the computational representation from which that answer emerged.

That gap suggests the need for a new object of evaluation. Current AI assurance largely measures outputs: factual accuracy, benchmark performance, hallucination rates, latency, security, robustness, and compliance. Those measurements remain essential, but they leave unanswered a prior question. How faithfully did the system reconstruct the reasoning presented by its human user before generating its response?

The answer should not depend upon intuition. Reconstruction can itself become computationally measurable. A system may preserve chronology accurately while altering evidentiary relationships. It may preserve individual facts while strengthening uncertainty into confidence. Attribution may remain attached to the correct speaker while scope quietly expands beyond what the original prompt supported. None of these failures necessarily produces an obviously false answer, yet each changes the reasoning inherited by the final response.

Such measurements would not determine whether a government reached the correct policy decision. They would answer a narrower and more fundamental question. Did the conversational system preserve the structure of human reasoning before assisting the official who relied upon it?

That question applies regardless of political system. The United States, Israel, China, Russia, NATO, and the European Union operate under sharply different political systems, legal authorities, military doctrines, alliances, and strategic aims. Each nevertheless faces the same practical question: what role should conversational artificial intelligence play in intelligence, military, diplomatic, legal, and medical decisions? The engineering challenge exists independently of the government deploying the technology.

International law already assumes that chronology, attribution, proportionality, available intelligence, and evidentiary relationships possess legal significance. An American targeting lawyer, an Israeli military legal adviser, a Russian intelligence officer, a Chinese security official, or an analyst working within NATO or an EU institution may operate under radically different legal and political systems, but each depends upon an accurate account of what was known, when it was known, and how confidently it was assessed. Those requirements do not disappear because information passes through a conversational model before reaching a human decision-maker. Widespread deployment therefore increases the importance of preserving those relationships rather than merely producing fluent prose.

Governments routinely preserve financial statements, classified access logs, procurement records, aircraft maintenance histories, laboratory procedures, election documentation, and chains of custody. Yet an agency may retain a prompt, the documents submitted to a model, and its final response without preserving any reviewable account of how chronology, confidence, attribution, or scope changed between them. An intelligence assessment, legal memorandum, clinical summary, or operational briefing can therefore enter the administrative record without exposing the computational reconstruction that shaped it. AI governance will remain incomplete until that intermediate transformation can be examined and measured.

Selected Sources

Brennan Center for Justice. 2025. “The Dangers of the Trump Administration’s Data Consolidation Efforts.” Brennan Center Research Reports, March 2025.

Carnegie Endowment for International Peace. 2026. “The Fog of AI War.” Strategic Europe, April 2026.

Chen, H. 2024. “Out of the Black Box: Uncertainty Quantification for LLMs via Conditional Probabilities.” Working Paper no. w34965. National Bureau of Economic Research.

Digital Constitutionalism Network. 2025. “Two Futures of AI Regulation under the Trump Administration.” Digi-Con Policy Analysis, January 2025.

European Union. 2026. “Regulatory Framework for AI: Guidelines on High-Risk Systems and Prohibited Practices.” European Commission Digital Strategy Policy Papers, June 2026.

Faris, L. 2024. “Algorithmic Targeting in the Iranian–Israeli Confrontation: Technical Realities, Legal Thresholds, and the Boundaries of Human Control.” F1000Research 14: 1200.

Fathallah, S. 2024. “Algorithmic Death-World: Artificial Intelligence and the Case of Palestine.” Public Humanities 1, no. 1.

Fayet, H. 2023. French Thinking on AI Integration and Interaction with Nuclear Command and Control, Force Structure, and Decision-Making. London: European Leadership Network.

Gao, Z. 2024. “Detecting and Evaluating Bias in Large Language Models: Concepts, Methods, and Challenges.” Journal of Behavioral Data Science 4, no. 1.

International Institute for Strategic Studies. 2026. “Military AI Governance under Strain: The US–China Dialogue.” IISS Online Analysis, June 2026.

Lawfare Institute. 2026. “The Missing Resistance in China’s AI Debate.” Lawfare Analysis, May 2026.

Parly, F. 2019. Artificial Intelligence in Support of Defence: Report of the AI Task Force. Paris: Ministère des Armées.

Powell, R. 2024. The EU AI Act: National Security Implications. London: Centre for Emerging Technology and Security (CETAS), Alan Turing Institute.

Sabaliauskaite, G. 2024. “Automated Regulatory Compliance for AI Systems in the Security Domain: The Case of Dual-Use Deployment.” Open Research Europe 4: 83.

Schmidt, E. 2021. Final Report: National Security Commission on Artificial Intelligence. Washington, DC: National Security Commission on Artificial Intelligence.

Sylvia, N. 2024. The Israeli Military’s Use of AI in Gaza: Operational Efficiency at the Cost of Humanity. Barcelona: IEMed.

U.S. Department of Defense. 2023. U.S. Department of Defense Responsible Artificial Intelligence Strategy and Implementation Pathway. Washington, DC: U.S. Department of Defense.

Vogiatzoglou, P. 2024. “The AI Act National Security Exception.” Verfassungsblog, May 2024.

Wang, Y. 2026. “Context Compression for LLM Agents: A Survey of Methods, Failure Modes, and Evaluation.” Preprints, May 2026.

Watts, T. F. A. 2025. “The Offset Imaginary: Great Power Competition, Security Imaginaries, and the Making of Artificial Intelligence in American Defense Planning.” Taylor & Francis Strategic Studies Journal.

Wiese, L., and C. Langer. 2024. “Gaza, Artificial Intelligence, and Kill Lists.” Verfassungsblog, April 2024.

Williams, T. 2025. “The Defence Economics of Artificial Intelligence & Machine Learning: Classification and Applications.” Taylor & Francis Defence and Peace Economics Journal.

Wolf, R. 2024. “Applying Precautions in Target Verification with AI Decision Support Systems.” Israel Law Review 57, no. 2.

Yang, D. 2026. “Emerging Patterns in Intelligentized Warfare: Speed, Consumption, and the Transformation of Military Operations.” Frontiers in Political Science 8: 1837867.

Yu, S. 2021. “Implications of AI in National Security: Understanding the Security Issues and Ethical Challenges.” Master’s thesis, Cardiff Metropolitan University.

About the Author
Kelsey Maurine Brickl is a historian and writer trained in Modern European History at the University of Edinburgh. Her work examines how truth is constructed, contested, and defended after mass violence, with a focus on Holocaust historiography, testimony, and archival evidence. She writes at the intersection of history, law, and public life, with particular attention to institutional accountability and disability rights.
Sign in or Register
Please use the following structure: example@domain.com
Or Continue with
By registering you agree to the terms and conditions
Register to continue
Or Continue with
Log in to continue
Sign in or Register
Or Continue with
check your email
Check your email
We sent an email to you at .
It has a link that will sign you in.