I'm always surprised to see that a large enough portion of the HN community are interested in classics that posts like this make it to the front page. Who are you all? What are you doing here?
My undergrad was in classics and viticulture, then did grad school in atmospheric science, which led me to tech and hackernews.
I think the answer is that the past is as relevant as the future. If people here know both computer programming and classical history, I think that means they have a more rounded view of the world.
The thing that strikes me from ancient texts is not the obvious differences from our day, but the similarities.
Interesting project. I would love if you could switch fonts to something like New Athena Unicode.
I built something similar to this by cloning the Diogenes repo and getting Claude to re-implement it in Python (it’s a very old battle tested Perl code base, so a great reference implementation) and using the TLG database for the Greek and Latin texts. You can take this even further by integrating it with the Barrington Atlas (there are scans on Anna’s Archive) for looking up ancient place names, so you have dictionary + map lookups. If you’ve read any of the Landmark series books you’d known what I mean.
Also better if you can generate chapter by chapter critical apparatus on difficult grammar and Anki decks. It’s an annoying part of learning these languages to have to stop and look up words, I like spending a few days learning vocab before reading and it makes it so much more pleasant than having to stop and look things up the whole time.
Obviously all this stuff is copyright so it can never be shared, but I don’t care it’s for my own personal use. I also bought like 6 hours worth of Ionnis Strattakis’ recordings where he reads a bunch of Ancient Greek in reconstructed Attic pronunciation and fine-tuned a text to speech model (styletts2) with full accent markings and breathings. It’s extremely natural sounding. My long term goal is to have a personal tutor that I can speak Attic to and basically have lessons with everyday (speech to text -> llm-> text to speech). All the pieces are there to actually do this.
LLMs have been an absolute game changer for me who is a hobby classicist but also knows how to build software. They are so amazingly good at Attic Greek and Latin, it’s like having the best teachers in the word at your fingertips for some really niche topics that it wouldn’t be possible with otherwise. Also extremely good at managing, building, cleaning up, deduplicating Anki decks.
> LLMs have been an absolute game changer for me who is a hobby classicist but also knows how to build software.
Definitely! I myself have been having a lot of fun[0] recently using LLMs to enrich one of my favorite beginner Latin readers[1] with audio forced-alignment, POS tagging, morphology, definitions, and other niceties that learners might appreciate.
Kinda font-related, it's throwing me off how every V in the Latin texts show as U. I don't like how they modernized the spellings in school, but if you're not doing that, aren't they all V?
Also most of the word popups use different spellings, like "iam" becomes "jam".
The way the Greek is displayed isn't very helpful, though. Specifically, any vowel with a grave accent displays the accent as a separate letter, which makes it very distracting to read. I don't think it's a problem with the encoding, as I can copy the inline text and it displays just fine - though with some weird spaces before commas or periods:
May I point out that in "Crudam si edes, in acetum intinguito", that "edes" is more likely to be the future of edere/esse "to eat"? (Just guessing by context.)
Suggestion: For the pop-up text, bold the meaning of the word so it pops out a bit better. Due to the formatting of word definitions, you have to go hunting for it in many cases.
Note that Perseus considers itself end-of-life; there is an analogous project that's supposed to have succeeded it, but I don't remember what project that is.
I built myself such a simple lexicon for technical stuff (concurrency vocab - invariant, genetics, stuff like that) - can only recom the practice as vocabulary is clearly a big step difference
I don't get a definition for 'fulgere', third word first entence, just a reference to 'fulgo'. I can guess what it means though from the more common 'fulgur', maybe something adjacent to flashing or lightning, but translations seem a bit sketchy.
That's actually ok. In virtually all Latin dictionaries verbs are listed in the first singular person of the present indicative (e.g. you would find "sum" and not "esse"). It's just showing you the dictionary entry.
It's not really very useful in this particular case, as you can't actually navigate the "dictionary" any other way than finding the base word form somewhere in the text you're looking at.
Clicking out to close a dictionary popup only works if you click empty space within the layout center. This is annoying. On a wide monitor, most empty space is on the sides (where clicking doesn't close).
It seems some books are missing chapter markings. Revelation, for instance. The verses are numbered, and you can notice when chapters change, but one can be lost when searching for a specific chapter.
Dictionary entries are cool, but I would want the in context meaning at least highlighted, so I don't have to read the full entry and do that myself.
Overall, nice idea but it seems like a very barebones implementation that needs a tremendous amount of polishing to be useful.
My undergrad was in classics and viticulture, then did grad school in atmospheric science, which led me to tech and hackernews.
How did you get here?
A lot of role playing games, especially from Square Enix,leverage a lot of Latin and classical education.
The thing that strikes me from ancient texts is not the obvious differences from our day, but the similarities.
I built something similar to this by cloning the Diogenes repo and getting Claude to re-implement it in Python (it’s a very old battle tested Perl code base, so a great reference implementation) and using the TLG database for the Greek and Latin texts. You can take this even further by integrating it with the Barrington Atlas (there are scans on Anna’s Archive) for looking up ancient place names, so you have dictionary + map lookups. If you’ve read any of the Landmark series books you’d known what I mean.
Also better if you can generate chapter by chapter critical apparatus on difficult grammar and Anki decks. It’s an annoying part of learning these languages to have to stop and look up words, I like spending a few days learning vocab before reading and it makes it so much more pleasant than having to stop and look things up the whole time.
Obviously all this stuff is copyright so it can never be shared, but I don’t care it’s for my own personal use. I also bought like 6 hours worth of Ionnis Strattakis’ recordings where he reads a bunch of Ancient Greek in reconstructed Attic pronunciation and fine-tuned a text to speech model (styletts2) with full accent markings and breathings. It’s extremely natural sounding. My long term goal is to have a personal tutor that I can speak Attic to and basically have lessons with everyday (speech to text -> llm-> text to speech). All the pieces are there to actually do this.
LLMs have been an absolute game changer for me who is a hobby classicist but also knows how to build software. They are so amazingly good at Attic Greek and Latin, it’s like having the best teachers in the word at your fingertips for some really niche topics that it wouldn’t be possible with otherwise. Also extremely good at managing, building, cleaning up, deduplicating Anki decks.
Definitely! I myself have been having a lot of fun[0] recently using LLMs to enrich one of my favorite beginner Latin readers[1] with audio forced-alignment, POS tagging, morphology, definitions, and other niceties that learners might appreciate.
[0]: https://hercules.hookbangsplat.com
[1]: https://archive.org/details/p1fablesoforbili00godl/mode/2up
Also most of the word popups use different spellings, like "iam" becomes "jam".
The way the Greek is displayed isn't very helpful, though. Specifically, any vowel with a grave accent displays the accent as a separate letter, which makes it very distracting to read. I don't think it's a problem with the encoding, as I can copy the inline text and it displays just fine - though with some weird spaces before commas or periods:
ΕΝ ΑΡΧΗ ἦν ὁ λόγος , καὶ ὁ λόγος ἦν πρὸς τὸν θεόν , καὶ θεὸς ἦν ὁ λόγος .
May I point out that in "Crudam si edes, in acetum intinguito", that "edes" is more likely to be the future of edere/esse "to eat"? (Just guessing by context.)
I don't get a definition for 'fulgere', third word first entence, just a reference to 'fulgo'. I can guess what it means though from the more common 'fulgur', maybe something adjacent to flashing or lightning, but translations seem a bit sketchy.
[1] https://latin-words.com
[2] https://latin-words.com/word/latin/fulgere
It seems some books are missing chapter markings. Revelation, for instance. The verses are numbered, and you can notice when chapters change, but one can be lost when searching for a specific chapter.
Dictionary entries are cool, but I would want the in context meaning at least highlighted, so I don't have to read the full entry and do that myself.
Overall, nice idea but it seems like a very barebones implementation that needs a tremendous amount of polishing to be useful.