Article URL: https://github.com/wolfpld/usenetarchive Comments URL: https://news.ycombinator.com/item?id=49081862 Points: 7 # Comments: 0

The Usenet Archive Toolkit project aims to provide a set of tools to process various sources of usenet messages into a coherent, searchable archive. People went away to various forums, facebooks and twitters and seem fine there. Meanwhile, the old discussions slowly rot away. Google groups is a sad, unusable joke. Archive.org dataset, at least with regard to polish usenet archives, is vastly incomplete. There is no easy way to get the data, browse it, or search it. So, maybe something needs to be done. How hard can it be anyway? (Not very: one month for a working prototype, another one for polish and bugfixing.) Why use UAT? Why not use existing solutions, like google groups, archives from archive.org or NNTP servers with long history? UAT provides a multitude of utilities, each specialized for its own task. You can find a brief description of each one below. Usenet messages may be retrieved from a number of different sources. Currently we support: Imported messages are stored in a per-message LZ4 compressed meta+payload database. Raw imported messages have to be processed to be of any use. We provide the following utilities: Raw data right after import is highly unfit for direct use. Messages are duplicated, there's spam. These utilities help clean it up: