A description of the TBX file structure—optimised for AI

The file was part of a quixotic and unlimitedly unsuccessful quest to get Claude to read and comment on the coherence of aTbRef. The AI can’t read the whole resource—at least not within the context available on a Claude Pro Plan I have no research funding to expand beyond that.

In discussion with Claude it started unequivocally that reading the raw XML was more useful/less wasteful than reading via MCP. Bear in mind we’re talking about reading the whole corpus not just a few note. It then became apparent that understanding general XML syntax is bot the same as understanding the format specific XML structure of a given type of file such as a TBX. So, using my existing aTbRef section (here) and a clean new TBX with at least one of everything that triggers any only-if-used XML data, I worked with Claude—deliberately in Chat mode not Code mode to make the above. Why Chat? This isn’t a coding problem even if the output is written in HTML. The hard graft was reverse-engineering an undocumented† XML structure and where some data is only present if used. IOW, it is easy to miss things neither I not the AI knew might be there. There are some benign edge cases—re reverse engineering—where the same-named element is used slightly differently in different place. No problem for Tinderbox as it knows the structure: less clear to parsing the structure without inside knowledge.

As well as the HTML page at thread start, the work also produced the TBX posted here.

Were a human or AI to actually read the HTML, the purpose is stated at start:

Purpose provenance and scope

Purpose
This document describes the XML structure of a Tinderbox (.tbx) file. It is intended to give an AI reader sufficient grounding to parse TBX content correctly, interpret XML attribute values accurately, and give well-informed advice to Tinderbox users about document structure and configuration. The syntax described is for version 2 of the TBX schema, as used in v6+ of the Tinderbox app.
Provenance
The document is based on information recorded in the resource "A Tinderbox Reference File" (aTbRef) at https://atbref.com. It is written by Mark Anderson, aTbRef's author in consultation with Tinderbox's designer Mark Bernstein (of Eastgate Systems). The document was generated using Tinderbox v11.7.1b805, at 2026-06-17T00:06:55+01:00.
Scope
The document does not cover Tinderbox's user interface, workflow, or conceptual design. Those are covered in the companion primer documents. This document covers the XML serialisation only.

Indeed, @Estomm’s Claude summary seems to pick this up. I’m slightly surprised at the its comment as to a ‘gap’ re JavaScript and AppleScript as these are self-evidently out of scope, so not fairly a weakness of the resource. Indeed, IIRC, Claude drafted the above so it can’t understand its own writing!

Documenting the JavaScript and AppleScript makes aTbRef’s scope look tiny. But anyone with expertise and lot of spare time for testing could fill that gap (my time is not so free at prresent).

I didn’t, it was the output of 2-3 man-days of effort spread over 2-3 weeks. If you factor in the aTbRef source material, it is years of experience in the mix.

This does seem an interesting use based on one of the key ideas—setting out information is an AI-friendly way that doesn’t waste tokens on aspects that are only desired by human readers.

As it happens I’d already realised this document alone didn’t resolve the original whole-aTbRef consumption problem because Tinderbox uses a link base so inter-note links aren’t harder-coded into notes as they with a wiki, for instance.

The resolution to the last can be seen at A Tinderbox Reference File. The epiphany was Claude observation that whilst there were c2.5 aTbRef HTML pages most of each pages was topic irrelevant (HTML header, body copy headers/footers). The method I used for older experiments to make PDF’s 'printed from a single HTML page hiply uses into one set of HTML stuctural code surrounding the body copy—the letter being the bit we want. After a quick sed call to change inter-HTML-page links to in-page links, we have a functional hypertext in a single HTML page.

This later lessens the context use (tokens) significantly. Even so, Claude’s (genAI’s) inability to sustain multiple parallel contexts—as human can with ease—still limits the AI’s ability to read the corpus. Still, it did result in Claude finding some inconsistencies whilst also throwing light on the shallowness of AI ‘Reading’. Most prompt questions tend to be over-reduced to one or more binary [sic] questions. So, fine for bounded factual errors but less use for broader qualitative.

I note the factual dstart error and have fixed it:
The_XML_TBX_format.html.zip (28.3 KB)

Apart from the above file, the most useful part of the experience was getting a clearer view of generative AI’s inabilities. They don’t devalue what the AI can do fast, accurately and all. But have a views as to its weaknesses has made it easier for me to avoid unhelpful output. I contingent challenge here for the human is genAI’s ability to generate a lot of text without obvious case errors, typos, etc, fools us into thinking it understands things in a human way. It does not!

†. I ought to add that the fact the TBX XML has no public documentation is not an error or omission. There has never been an undertaking to offer that. My older aTbRef notes were a simple fade mecum for the few users who like to, or occasionally needed to, peeked under the hood at file info. It was only Claude insistence that XML was more use than in-app MCP or HTML as an input source in the whole doc/website sense that made me realised so documentation might be useful. The ‘schema’ is my work and any mistakes are mine and Claude’s, not Eastgate’s.

1 Like