That was a fascinating read, I really enjoyed it. Literally like exploring lost knowledge. Great work and a great write-up.
I also liked the aesthetics of it and the little effects (meteorite and volcano, but please fix the rhino and the text flowing around it while it rotates).
I wonder what else could be found in such archives. Some ideas:
- Locations or routes of sunken ships and their missing cargo?
- Some pirate stories, maybe about a now-forgotten but once-legendary pirate captain?
- Unusual weather events, like snow in the summer?
IMHO: The rotating rhino, meteor impact, and animated flowchart is totally unnecessary cruft that makes it look almost satirical. If this keeps up, in time, this "AAA effects" stuff is going to look like the 90s "under construction" banner gifs.
The effects are comically bad. I see the inspiration in scrolling effects that the New York Times put together, but the NYT was never dumb enough to obscure the copy text. Form follows function, and the function of a web page is to be read, not to obscure what is to be read with some stupid effect that's supposed to remind one (I suppose) of a volcano's cloud obscuring one's vision. At least "under construction" banners didn't obtrude upon the copy text.
>To make this kind of research accessible, I’m open-sourcing the workflow I created for this investigation as a small toolkit, Antiquity, enabling anyone with a question and a coding agent to conduct similar historical archival investigations.
There are also VOC archives at Cape Town, also in Kew (search for the letters of Loot) which were literally looted by privateers. All these are written in High Dutch some in German. How reliable are the translations?
I'm working on a similar project for contemporary political opinion media. Every podcast, blog, oped, or show cut into little pieces with the structure, speaker, quotes and nouns pulled out and cross-referenced. I bring it up because I wonder if this kind of heavy-weight preprocessing is worth bringing to historical documents as well. It would be much more expensive, initially, but afterwards allows questions get answered even cheaper than they are in your current system. It may be worth collecting interested parties and co-investing in the structured parsing.
Also modern transcription and historical document scanning have a similar shaped problem - dealing with misspelled words and trying to infer their corrections from context.
Please just make sure to keep the ethical implications of any such work in mind.
I do not know what exactly it is you're building, but the shape also fits "weapon", and weapons do not really care about the good intentions of their author.
>Also modern transcription and historical document scanning have a similar shaped problem - dealing with misspelled words and trying to infer their corrections from context. Very true in my case on similar problems, my major issue was OCR relics. Reasonable mispelled words say by an uneducated person, are not that much of an issue. For the OP VOC work most letters were written by educated scribes and less of a problem. Anything before 1650 had very
different calligraphy though.
If I sat down to read just the Dutch East India Company pages myself, at two minutes a page, eight hours a day, five days a week, it would take me about 70 years. And that’s before the newspapers. My homebrew AI lab got through the entire archive in a single twelve-hour overnight run.
Makes me wonder how much the author himself learned about the Dutch East India Company. I suspect very little, if anything. Something about these exercises reminds me of junk food: empty calories and all that...
You know you don't have to comment on everything you see on the internet, right? You're allowed to just keep scrolling when something isn't for you. I wonder how much joy you have in your life, I suspect very little, if anything.
I didn't mean to upset you, piratebroadcast. Like so many of us, you scratched a technological itch, and that is fine.
However, now that AI tools have supercharged us, these scratches are easier to scratch, and their outcomes, made public, flood the aether.
Specifically that these rabbit holes are useful to bring people to all sorts of new discoveries and skills.
It is important to point out that that's a real risk with such AI use.
Of course, it is also true that it would likely not have happened at all otherwise. Both things can be true at the same time.
__
Oh I just realized that you're the actual author and this is not the only super-thin-skinned comment.
Possibly, yeah, but the reactions here do not look like the learnings have been long-term-stable building blocks, tbh.
But anyway.
I am curious what else will show back up again when other people decide to hook up a clanker to one of the many piles of historic (and current) data we do not have the manpower to process for.
I honestly primarily enjoy my browser performance not tanking on this $5000 workstation and actually bailed out while scrolling because it wasn't really usable.
I also liked the aesthetics of it and the little effects (meteorite and volcano, but please fix the rhino and the text flowing around it while it rotates).
I wonder what else could be found in such archives. Some ideas: - Locations or routes of sunken ships and their missing cargo? - Some pirate stories, maybe about a now-forgotten but once-legendary pirate captain? - Unusual weather events, like snow in the summer?
Just calling the progression while we're on it.
https://github.com/jessewaites/antiquity
I'm working on a similar project for contemporary political opinion media. Every podcast, blog, oped, or show cut into little pieces with the structure, speaker, quotes and nouns pulled out and cross-referenced. I bring it up because I wonder if this kind of heavy-weight preprocessing is worth bringing to historical documents as well. It would be much more expensive, initially, but afterwards allows questions get answered even cheaper than they are in your current system. It may be worth collecting interested parties and co-investing in the structured parsing.
Also modern transcription and historical document scanning have a similar shaped problem - dealing with misspelled words and trying to infer their corrections from context.
I do not know what exactly it is you're building, but the shape also fits "weapon", and weapons do not really care about the good intentions of their author.
Makes me wonder how much the author himself learned about the Dutch East India Company. I suspect very little, if anything. Something about these exercises reminds me of junk food: empty calories and all that...
Specifically that these rabbit holes are useful to bring people to all sorts of new discoveries and skills.
It is important to point out that that's a real risk with such AI use. Of course, it is also true that it would likely not have happened at all otherwise. Both things can be true at the same time.
__
Oh I just realized that you're the actual author and this is not the only super-thin-skinned comment.
Man. Why do be like this.
But anyway.
I am curious what else will show back up again when other people decide to hook up a clanker to one of the many piles of historic (and current) data we do not have the manpower to process for.
Using Opus 5.5 to discover a new eyewitness record of the dodo - https://news.ycombinator.com/item?id=49926917 - Oct 2026 (79 comments)
I bet it works great in chrome tho