Essay · Touhou reconstruction

The industrial revolution in reverse engineering.

By N0zoM1z0 · 10 October 2026

Takeaway

  • Give agents autonomy over the investigation. Set a clear goal and acceptance criteria, then let the agent decide what to try next.
  • Use the repository and Git as working memory. Keep the reasoning with the code. Git checkpoints let the next session continue after the conversation ends.
  • Let evidence decide what you claim. A confident explanation still needs testing. “Unknown” is a valid answer when the evidence is incomplete.
  • Test the oracle too. Try cases it should reject. If it misses a bug, revisit the earlier results that depended on it.
  • We are in the early days of RE's industrial revolution. The methods are still taking shape. There is plenty left to explore.

How experience compounds

Human direction · Goal and acceptance criteria

Autonomy

Agent

Choose an experiment.
Propose a hypothesis.

Verification

Oracle

Test against a
concrete reference.

Repository

Working memory

Code, evidence,
checks and lessons.

Mismatch
Revise the hypothesis

A better
starting point

Reuse

Pass · Keep the result and its evidence

Each project leaves the next with a better starting point. That includes fixing the checks when we discover a gap.

What changed in August 2026

Before this work, I had done a lot of manual reverse engineering. I would inspect a function until I had an explanation, then test that explanation against the program. Getting further meant spending more of my own time on the next function. It also meant keeping an increasingly large picture in my head.

Most of my RE work is on games. I reconstruct them so their behavior can be understood and eventually carried into a port or a mod. That is the experience behind this article. The method can be useful elsewhere, but each field needs to establish what it can check against a reliable reference.

In August 2026, I started exploring agent-driven RE seriously. I wanted to see how far an agent could get with access to the program and ordinary development tools. Touhou reconstruction gave me a substantial project in which to find out.

The early progress was startling. An investigation could keep moving without me choosing every individual step. A failed comparison might send the agent to another caller. It could write a small diagnostic and use the result to decide what to try next. Giving it that freedom made a difference: I could spend more attention on the project's direction and on whether its checks were trustworthy.

The pace got my attention first. Then I began wondering what would survive beyond the current project. Would the next game benefit from everything we had just learned?

TH08: building on human work

TH08, Imperishable Night, began as a continuation of GensokyoClub's reconstruction. Their public source gave me a substantial foundation. It came with knowledge of the build and a history of contributions that I preserved in the continuation.

The imported history ends at the public 10 August checkpoint. My independent continuation began on 13 August. By 19 August, the ledger recorded source for all 1,107 identified game functions.

A playable Linux reconstruction port was committed on 24 August, roughly eleven days after the continuation began. A Web edition followed. The native Linux 64-bit release arrived on 30 August.

What struck me was that the work had reached a program people could run. Getting there required following problems beyond an individual function. Even when two functions looked correct in isolation, they could accidentally use separate copies of state that should have been shared.

This made reference checks central to the work. I call them oracles. A code comparison can tell us whether a reconstructed function reproduces the original instructions. A runtime check can tell us whether an exercised route reaches the expected state. The agent proposes an explanation and tests it; a mismatch gives it something specific to investigate.

An exact match with the wrong constant

An oracle is software someone wrote. TH08 gave me a clear reminder of how much can depend on that software.

In September, a bug reported by a downstream Switch port led back to item auto-collection. The reconstructed source checked player power against 0.0. The original used 128.0, the maximum-power threshold.

Normal power is non-negative, so the reconstructed power check was effectively always satisfied. Above the collection line, the game could attract items without requiring maximum power. The other exceptions in the condition were correct. This one constant changed the behavior.

Yet the function had already passed the exact comparison.

A rebuild can put a constant at a different address. The comparison tool accounted for that by adjusting addresses in the compiled instructions before comparing them with the original. But it never checked the floating-point value stored at the referenced address.

That left a hole in the check. It could point the reconstructed instruction at the original's 128.0 while the source still said 0.0. The instruction bytes matched. The source meant something different.

The 2 September fix corrected the source and made the comparison check the referenced constant's actual bytes. A wider audit then examined 1,548 floating-literal references. It found another twelve incorrect references across five accepted functions.

The comparison now checks every such floating literal. Its tests include deliberately wrong values to make sure they are rejected. We also had to revisit results that had passed the weaker check.

The oracle needed reconstructing too. Repairing it was part of repairing the game.

That matters when the team is mostly one person working with agents. I cannot personally inspect every line they produce, so much of my confidence rests on their checks. A shared oracle's blind spot can affect many investigations before I notice it. I have to understand what the checker actually verifies and test it against cases that ought to fail.

As the work continued, those checks became part of the repository. So did the reasons for source changes and notes that let a later session pick up where an earlier one stopped. The repository was becoming the project's working memory.

A successful batch gave us both recovered code and a better environment for the next batch.

Two different ideas of reconstruction

GensokyoClub's public README makes its disagreement with this kind of work explicit. The notice announces that further development will happen privately until completion. One passage reads:

The rise of grifters (AI decompilations and ports) in this space taking from our work paints a bad image for future decompilation efforts…

The notice also describes the psychological toll on the maintainers. Their contribution policy excludes pull requests produced primarily with AI. They have put their free time into difficult work, and that work helped make my continuation possible. I respect the effort behind it. The disagreement I want to discuss is about how reconstruction should proceed and how a contribution should be judged.

In the manual workflow I knew, developing an understanding of a function and reconstructing it usually fell to the same person. A project relied heavily on that person's expertise. Trust in the contributor mattered because so much of the reasoning happened while they worked.

Existing projects already preserve knowledge in their source and build tools. What changed for me was that an agent could use that knowledge to pursue the next investigation on its own.

In my continuation, I decide the objective and the standard for accepting a result. The agent has broad freedom to investigate. Its proposed reconstruction then has to survive the relevant checks. I want another person to be able to examine why we chose an implementation, even if an agent did most of the work.

This can be a difficult transition. Years of careful work may become the foundation for a continuation that moves much faster. That raises real questions about credit. It also changes what maintainers need to know before accepting a contribution.

The industrial analogy helps me think about this. In a craft, much of the process depends on the skill of the person carrying it out. Machinery changes where that skill is needed. Someone still has to design the process and recognize when its output is wrong. Different communities can choose how much of that change they want to take on.

My choice is to continue openly, with the inherited work credited and its history preserved. I want the new work to be reviewable. That gives us a way to ask how far this approach can go and to learn from what goes wrong along the way.

Source: GensokyoClub's README notice, checked on 10 October 2026, and its contribution policy. The quotation is a shortened excerpt. TH08's credits and provenance record the continuation boundary.

TH095: experience starts compounding

TH095, Shoot the Bullet, made the value of that experience much easier to see. Its repository began on 29 August with zero confirmed game functions in the ledger. We still had to learn the game. But we already knew much more about how to begin a reconstruction and how to keep it moving.

By 7 September, all 697 identified game functions had source. By 8 September, 696 had been accepted as exact comparisons. The whole program linked on 9 September. The Windows i386 reconstruction was marked playable on 10 September, about twelve days after initialization.

I found this more exciting than the first project's speed. A fresh target could benefit from work done on another game. The experience was already present in the tools and in the way the project was organized.

For example, TH08 had taught us to pay attention to the assembled program early. If several recovered functions depend on the same state, their isolated comparisons leave an important question open. We need to see them working together. That lesson helped shape how we approached TH095's whole-program build.

A lesson that stays in my head is useful while I am there. Once it becomes a check that another session can run, it can keep helping after I have moved on. The next agent can use the result without repeating the investigation that led to it.

The floating-literal bug belongs in that memory too. It explains why checking a reference also requires checking the data behind it. Keeping that correction with the code helps later projects avoid inheriting the old check's blind spot.

The method becomes part of the starting material for the next game. We can spend more of the next project's effort on what is actually new about its target.

The same benefit is available to people joining later. They can inspect a decision and rerun its check before continuing the work. They do not have to reconstruct the entire project's history to find out why the source looks the way it does.

TH04: the workflow survives a different architecture

TH04, Lotus Land Story, took this work into the PC-98 DOS era. The target was now a 16-bit environment with four cooperating programs. Understanding its hardware behavior required different evidence from the Windows games. Existing ReC98 work gave us valuable knowledge and source material here too.

The DOS reconstruction is now working. In my manual testing, I played complete Normal routes through their endings and checked saves. The current work is a native 64-bit port. Establishing a working DOS version first gives that port a reference.

The architecture changed what we needed to investigate. It also changed the compiler and runtime against which we checked our work. But the agent could still follow a question through to a result and let that result guide the next experiment.

Consider the transition from gameplay to an ending. We need to know what state carries across that boundary and which program is responsible for it. That is something we can investigate against the DOS product. Once the evidence and checks are available, an agent can work through the question much as it did on a Windows title.

This is why TH04 matters to the argument. A substantial platform change did not force us to start over with a new way of working. The architecture defined the problem; the workflow still gave us a way to solve it.

For the 64-bit port, we can now examine the new implementation against behavior already recovered on DOS. The knowledge from reconstruction gives the port something to build on.

Project state as of 10 October 2026: DOS reconstruction and manual testing · 64-bit port. The port remains in development.

From exact code to readable code

Once a reconstruction works, I want someone else to be able to understand it.

To me, assembly and raw offsets can feel like old friends. I realise this is a slightly unusual definition of “readable.” Most people would prefer to follow the game's logic without keeping the executable's memory layout in their head.

That is where semantic reconstruction comes in. A recovered field may still be known mainly by its offset. We follow how the game uses it until we can explain its role. Then we can give it a meaningful name and a type that fits the evidence. We keep the reasoning with the source so the next person can see where that interpretation came from.

This becomes especially important for a port. An absolute address tells me where something lived in the old executable. It gives a 64-bit implementation little help in deciding which object should own that state. To move the behavior safely, we need to recover the relationship behind the old memory access.

The order I now use is:

  1. Recover an exact baseline. Compile the reconstructed pieces with the historical compiler. Compare the relevant code and data with the original executable. Record unresolved differences so the next phase has a clear starting point.
  2. Build and play it on the original platform. Link those pieces into the real program using the original architecture and compiler. Exercise important gameplay paths. This is where we can find problems with shared state or initialization that an isolated function comparison missed.
  3. Reconstruct the semantics against both references. Take one coherent part of the game at a time and establish what its recovered source means. Improve its representation while preserving the exact comparisons and the playable historical build.
  4. Make the modern port. Move the established behavior into the new environment, such as a native 64-bit build. The reconstructed original-platform game remains a reference for comparing how the port behaves.

The playable build from the second stage becomes a second oracle during the third. The first oracle checks whether our changed source still reproduces the relevant original code and data. The second checks that the reconstructed program still builds and behaves correctly along the paths we exercise.

They catch different mistakes. A type change can alter the generated instructions. An ownership change can leave two parts of the game using different copies of state. Keeping both checks available gives the agent a concrete failure to investigate before carrying a refactor further.

A name needs evidence of its own. An exact comparison cannot tell us whether a field really means “invulnerability time.” We have to establish that from how the game writes and uses it. If the meaning remains uncertain, a neutral name is more helpful to the next reader than a confident guess.

We learned this order through the projects. TH08 already had playable ports before some of its later historical-platform audits. That made certain defects harder to see. The Factory's current workflow puts the original-platform build first, so semantic work can use it as a reference before porting begins.

Exact reconstruction gives us a reference. Semantic reconstruction makes the recovered knowledge usable. A port can then build on both.

What makes this an industrial shift

These projects changed where I spent my attention. Once agents could carry much of an investigation forward, improving their working environment became one of the most useful things I could do. A better tool could help with every later function that needed it.

Autonomy matters here. The useful next step often becomes clear only after a failed experiment. An agent needs enough freedom to follow that result somewhere unexpected. If it has to wait for me to prescribe each step, much of the work remains tied to my attention.

I expect the agent to make wrong hypotheses. What matters is whether we can test them and learn from the result. A failed check should help it understand the mistake well enough to try again. I still have to decide whether the accumulated evidence supports a project milestone.

REA gives the agent access to analysis tools. A question about a function's caller can lead straight into inspecting that caller. The reconstruction project supplies the compiler and its own reference checks. The agent can use those to test the source it proposes and see where its explanation holds up.

The TH08 bug shows why those checks deserve engineering attention of their own. When the same comparison is used across hundreds of functions, a gap in it can spread much further than a mistake in one implementation. Testing the checker improves the feedback available to all that later work.

The industrial analogy has a useful historical example here. Boulton and Watt introduced a steam-engine indicator in 1796 to help adjust engine valves. A recording version traced pressure through the piston stroke. It made the engine's internal behavior available for inspection. Our comparison tools serve a similar purpose: they let us examine what the machinery is doing while we improve it.

We are in the early stages of this industrial shift. Much of the infrastructure is still immature. Agents can move faster than our checks were designed to support, so the process has to develop alongside them. When we find a defect in a shared tool, we have to repair it and revisit the results it affected. The next project can then inherit a stronger tool.

There is also a practical limit to any one conversation. It will end before a large reconstruction is finished. The repository has to make it possible for another session to continue without losing the reason for the last decision.

The Touhou Reconstruction Factory grew out of that need. It gives projects a shared way to carry their checks and lessons forward. Work on one game can improve the starting conditions for another.

That is what makes the industrial analogy meaningful to me. Experience starts becoming part of tools that others can use. Improving those tools changes how much the next person—or the next agent—can get done.

The projects we can now consider

The pace matters because it changes the decision to begin. A game could be fascinating to reverse engineer and still demand more of my own attention than I could realistically give it. Many projects would stay ideas.

Now I can see a way to keep such a project moving through repeated investigations. Getting a working reconstruction makes a port more practical. Recovering readable semantics makes it easier for someone else to explore a mod. The effort put into understanding the game can keep paying off after the first version runs.

I now look at an unfamiliar program and ask: what access, feedback and accumulated knowledge would let an agent work on this reliably?

That question makes me consider projects I would previously have left alone. Each one can improve the way we approach the next. I want to keep exploring how far that can take us.

Project milestones and sources

The dates describe recorded project checkpoints, checked against public GitHub history on 10 October 2026. Elapsed time is calendar time between commits. Source presence, exact comparisons, a build and a runtime result each name a different milestone.

More articles · Explore a TH04 investigation

Top