Short version of the situation is that I have an old site I frequent for user written stories. The site is ancient (think early 2000’s), and has terrible tools for sorting and searching the stories. Half of the time, stories disappear from author profiles. Thousands of stories and you can only sort by top, new, and 30-day top.

I’m in the process of programming a scraper tool so I can archive the stories and give myself a library to better find forgotten stories on the site. I’ll be storing tags, dates, authors, etc, as well as the full body of the text.

Concerning the data, there are a few thousand stories- ascii only, and various data points for each story with the body of many stores reaching several pages long.

Currently, I’m using Python to compile the data and would like to know what storage solution is ideal for my situation. I have a little familiarity with SQL, json, and yaml, but not enough to know what might be best. I am also open to any other solutions that work well with Python.

  • amenji@programming.dev
    link
    fedilink
    arrow-up
    2
    ·
    7 months ago

    A lot of people already suggests several databases or plaintexts like json.

    But to be honest if the dataset is not too big and doesn’t grow (since it is historical anyway), why not just use markdown with Hugo (a static site generator). You could also make use of its supported search tools to search texts in the stories.

    As a bonus, since it’s a static website, you can host it and share it to the world!

    • Bubs@lemm.eeOP
      link
      fedilink
      arrow-up
      2
      ·
      7 months ago

      I’ll give it a look. I’m still in the early stages of the project, so it’ll be a bit before I get to the point where I work on the database side of things.