important notice: this article doesn’t describe anything that is at all useful in any way. this is interesting and cool but NOT useful
accidental poetry
a few friends of mine occasionally write poems. a couple of them write on pen and paper and take their poetry pretty seriously and even publish some of it, while most others i know mainly write on their cell phone. my best friend Benedetta is amongst the cell phone people, and a few months ago she asked me if i could help back up her 60 or so poems from the proprietary third party note-taking android app she was using into some other safer format for long-term storage. i tried to, failed, decided to attempt the task later, and then forgot to do it. i don’t have the backup file anymore, but i remember opening in a hex editor showed lines like com.app-name.ClassName or something similar, which made me think it was some sort of Java archive i could have easily extracted the data with some Java-based tools. as a sort of quick attempt and proof of concept, i tried to extract text from the binary archive with Unix’s (my Macbook, specifically) strings utility. very interesting stuff came out.
from my content of the day in january:
2026-01-05
more mangled poems:
certi giorni mi rispondo
che non passer
mai.sq
0Lxsq
+sono ancora il tuo fiore?
2026-01-05
i tried to help my friend Benedetta backup her notes app with all her poems and trying to open the file with strings (unix utility) it came out a little broken and i really enjoyed the way it broke:

an interesting thing about this is that apparently the strings are not encoded with UTF-8 or at least something weird happens with non-ASCII characters (the Italian à, è, é, ò, ì, ù). from a couple of lazy and imprecise experiments i think that’s because Apple’s strings program considers any byte >0x80 to be a “non printable character” so the two sections before and after a special character are counted as two different text fragments. but i’m not too sure.
i was imagining myself throughout this experiment as performing some sort of surgery on this data to recover it from a machine that took it against a human’s will. that’s a very silly way of looking at it: we gave it the poems for free
sqlite
i tried to recreate this accidental datamoshing on purpose, with my friends’ poems from the Femministe nella Foresta collective.

i scraped their public poems with a python script and the Wordpress API:
def parse_post(post: dict) -> dict:
raw = post["content"]["rendered"]
# Author: first <p class="wp-block-paragraph">Name</p> block before the
# verse (sometimes wrapped in <strong>, sometimes plain text).
author_match = re.search(
r'<p class="wp-block-paragraph">(.*?)</p>', raw, re.DOTALL
)
author = html_to_text(author_match.group(1)) if author_match else "Sconosciuto"
author = author or "Sconosciuto"
# Body: everything inside <pre class="wp-block-verse">...</pre>
body_match = re.search(
r'<pre class="wp-block-verse">(.*?)</pre>', raw, re.DOTALL
)
body = html_to_text(body_match.group(1)) if body_match else ""
return {
"id": post["id"],
"slug": post["slug"],
"title": html_to_text(post["title"]["rendered"]),
"author": author,
"date": post["date"],
"url": post["link"],
"category_ids": post["categories"],
"body": body,
}
put them in a sqlite database:
import json
import sqlite3
import sys
import urllib.parse
from pathlib import Path
BASE_DIR = Path(__file__).parent
POEMS_DIR = BASE_DIR / "poems-01"
METADATA_FILE = BASE_DIR / "poems-metadata.json"
SCHEMA_SQL = """
CREATE TABLE IF NOT EXISTS authors (
id INTEGER PRIMARY KEY AUTOINCREMENT,
name TEXT NOT NULL UNIQUE
);
CREATE TABLE IF NOT EXISTS categories (
id INTEGER PRIMARY KEY AUTOINCREMENT,
name TEXT NOT NULL UNIQUE
);
CREATE TABLE IF NOT EXISTS poems (
id INTEGER PRIMARY KEY AUTOINCREMENT,
wp_id INTEGER UNIQUE, -- original WordPress post id
slug TEXT NOT NULL UNIQUE,
title TEXT NOT NULL,
author_id INTEGER NOT NULL REFERENCES authors(id),
date_published TEXT, -- ISO 8601, as published on the site
source_url TEXT,
body TEXT NOT NULL, -- full poem text
txt_file TEXT NOT NULL -- filename within poems-01/
);
CREATE TABLE IF NOT EXISTS poem_categories (
poem_id INTEGER NOT NULL REFERENCES poems(id),
category_id INTEGER NOT NULL REFERENCES categories(id),
PRIMARY KEY (poem_id, category_id)
);
CREATE INDEX IF NOT EXISTS idx_poems_author ON poems(author_id);
CREATE INDEX IF NOT EXISTS idx_poem_categories_category ON poem_categories(category_id);
"""
def get_or_create(cur, table, name):
cur.execute(f"SELECT id FROM {table} WHERE name = ?", (name,))
row = cur.fetchone()
if row:
return row[0]
cur.execute(f"INSERT INTO {table} (name) VALUES (?)", (name,))
return cur.lastrowid
def main():
out_path = Path(sys.argv[1]) if len(sys.argv) > 1 else BASE_DIR / "db-01.db"
poems = json.loads(METADATA_FILE.read_text(encoding="utf-8"))
if out_path.exists():
out_path.unlink()
conn = sqlite3.connect(out_path)
cur = conn.cursor()
cur.executescript(SCHEMA_SQL)
for poem in poems:
author_id = get_or_create(cur, "authors", poem["author"])
readable_slug = urllib.parse.unquote(poem["slug"])
txt_file = f"{readable_slug}.txt"
cur.execute(
"""INSERT INTO poems
(wp_id, slug, title, author_id, date_published, source_url, body, txt_file)
VALUES (?, ?, ?, ?, ?, ?, ?, ?)""",
(
poem["id"],
poem["slug"],
poem["title"],
author_id,
poem["date"],
poem["url"],
poem["body"],
txt_file,
),
)
poem_id = cur.lastrowid
for cat_name in poem["categories"]:
cat_id = get_or_create(cur, "categories", cat_name)
cur.execute(
"INSERT OR IGNORE INTO poem_categories (poem_id, category_id) VALUES (?, ?)",
(poem_id, cat_id),
)
conn.commit()
n_poems = cur.execute("SELECT COUNT(*) FROM poems").fetchone()[0]
n_authors = cur.execute("SELECT COUNT(*) FROM authors").fetchone()[0]
n_cats = cur.execute("SELECT COUNT(*) FROM categories").fetchone()[0]
conn.close()
print(f"Wrote {out_path}")
print(f" poems: {n_poems}")
print(f" authors: {n_authors}")
print(f" categories: {n_cats}")
if __name__ == "__main__":
main()
went into the sqlite file and manually removed a column (the idea is to create some confusion and fragmentation in the file, although i don’t know how sqlite works under the hood well enough to tell if this has any effect on the end result). ideally i would want to build my script so that it also occasionally deletes some records and edits others, adds junk in between useful records and removes it at the end, and other similar stuff to mangle the end result even more.
$ sqlite3 db-01.db
sqlite>
sqlite>
sqlite> ALTER TABLE poems DROP COLUMN source_url;
sqlite>
sqlite> SELECT name, txt_file FROM poems JOIN authors ON poems.author_id = authors.id;
Mauro Fusar Poli|prima-e-dopo-di-te.txt
Mauro Fusar Poli|uscire-di-casa.txt
Gaia Incarbone|adronitis.txt
Gaia Incarbone|pompelmo.txt
[...]
sqlite> .exit
the results could be more interesting.
this is the poem Eterno by Mauro Fusar Poli:
Ciò che ci siamo detti in questa vita
guiderà i nostri spiriti nella prossima,
cosicché il tempo venga ingannato
e potremo innamorarci ancora.
and this is how it comes out of the process:
eternoEterno
2025-05-21T09:10:58Ci
che ci siamo detti in questa vita
guider
i nostri spiriti nella prossima,
cosicch
il tempo venga ingannato
e potremo innamorarci ancora.eterno.txt
update: pebbledb poetry
the storage engine is the part of the database management engine that handles storing the records and indices on disk. i figured if i would try this experiment with a different storage engine i would get different and hopefully more interesting results. i tried with pebbledb, the storage engine at the heart of cockroachdb.
i used go this time since it’s the language pebble is written in.
here’s a script that:
- stores all the poems one-by-one in order multiple times
- deletes all the poems in a random order until the database is empty
- repeats the process multiple times
- stores all the poems one-by-one one final time
package main
import (
"fmt"
"math/rand"
"os"
"path/filepath"
"strings"
"uuid"
"github.com/cockroachdb/pebble"
)
const (
poemsDir = "poems-01"
dbPath = "poems.data"
writesPerRound = 20
rounds = 10
)
type poem struct {
name string
content []byte
}
func main() {
if err := run(); err != nil {
fmt.Fprintf(os.Stderr, "error: %v\n", err)
os.Exit(1)
}
}
func run() error {
poems, err := loadPoems(poemsDir)
if err != nil {
return err
}
db, err := pebble.Open(dbPath, &pebble.Options{Logger: nil})
if err != nil {
return fmt.Errorf("opening pebble db %s: %w", dbPath, err)
}
defer db.Close()
wo := &pebble.WriteOptions{}
for round := 1; round <= rounds; round++ {
keys, err := writeAllNTimes(db, poems, writesPerRound, wo)
if err != nil {
return fmt.Errorf("round %d: writing: %w", round, err)
}
if err := deleteAllRandomOrder(db, keys, wo); err != nil {
return fmt.Errorf("round %d: deleting: %w", round, err)
}
}
if _, err := writeAllNTimes(db, poems, 1, wo); err != nil {
return fmt.Errorf("final write: %w", err)
}
return nil
}
func loadPoems(dir string) ([]poem, error) {
entries, err := os.ReadDir(dir)
if err != nil {
return nil, fmt.Errorf("reading %s: %w", dir, err)
}
poems := make([]poem, 0, len(entries))
for _, entry := range entries {
if entry.IsDir() {
continue
}
path := filepath.Join(dir, entry.Name())
content, err := os.ReadFile(path)
if err != nil {
return nil, fmt.Errorf("reading %s: %w", path, err)
}
_, body, ok := strings.Cut(string(content), "\n\n")
if !ok {
return nil, fmt.Errorf("%s: no blank-line separator found (expected header\\n\\npoem)", path)
}
poems = append(poems, poem{name: entry.Name(), content: []byte(body)})
}
return poems, nil
}
func writeAllNTimes(db *pebble.DB, poems []poem, n int, wo *pebble.WriteOptions) ([]string, error) {
keys := make([]string, 0, len(poems)*n)
for i := 0; i < n; i++ {
for _, p := range poems {
key := uuid.NewV4().String()
if err := db.Set([]byte(key), p.content, wo); err != nil {
return nil, fmt.Errorf("writing key %q (%s): %w", key, p.name, err)
}
keys = append(keys, key)
}
}
return keys, nil
}
func deleteAllRandomOrder(db *pebble.DB, keys []string, wo *pebble.WriteOptions) error {
for _, idx := range rand.Perm(len(keys)) {
key := keys[idx]
if err := db.Delete([]byte(key), wo); err != nil {
return fmt.Errorf("deleting key %q: %w", key, err)
}
}
return nil
}
i think writing all the poems in the same order every time creates a more coherent result than shuffling them.
here’s Gaia’s L’uovo:
764914-e86
-Xab-8852-ade35e8b78b7
73797-6275-4776-9fa9-ea02529ddb58
db7e5-17e3-4f31-a442-72db8ada4521
8790be-d83f-4d58-8ab6-03d398bad4a6
Io non sono un albero, ma lo ero;
!@ho pi
le foglie
$(e avevo.
X nel tronco dell
eban
ro;
ram
($ faggio co
stata
+8e poi trave;
mollusc
l gremb
mare.ZI
cielo di
\alve
Ho letto ogni lib!$
ascolta
0canzone;
parl
4lingua, indoss
4colore;
sceso,
milione
volte,!k
scale1Jandr
\morso ed ingoiato, vivo,
sapo!
)Hmadr%
, dolce
ama)
!bPquel mio figlio mai n
a<artista e scienz
E, come Uliss
VDesidero essere ign
B8poeta. L
aman
to.a
d42a0-d7a0-4872-b6be-364efb69fe57
5c54-db90-4099-978e-c6360f99278c
f2a1-6b52-4a32-93ad-8ed877eaab5e
95ae18-6f50-41da-af77-680fc127020d
here’s a shell command to remove the UUIDs sprinkled around (ex: 73797-6275-4776-9fa9-ea02529ddb58)
$ strings poems.data/000013.sst | grep -vE '^[a-f0-9\-]+$' > poems.txt
P.S.
GNU binutils’ strings does have the -e flag to select character width and the mac has the iconv utility to filter only valid UTF-8.
/opt/homebrew/opt/binutils/bin/gstrings -e S db-01.db | iconv -f utf-8 -t utf-8 -c > strings-01.txt
but valid UTF-8 is boring for this purpose. also, i don’t have binutils in my $PATH