notebook: add docx→markdown converter for openwebui knowledge bases
- 000-configs/notebooks/docx_to_markdown_openwebui.ipynb (29 cells) - mammoth + python-docx fallback, manifest dedup (SHA-256), pre-flight fuzzy dedup (Jaccard @ 0.97), URL extraction, gzip download - 000-configs/notebooks/README.md: index of notebooks + how-to-add - 000-configs/README.md: list notebooks/ subfolder in the table - 07-tasks/docx-to-markdown-converter-notebook.md: full task writeup
This commit is contained in:
@@ -21,6 +21,7 @@ Server, model, and workspace configuration docs. Mirror of `/home/user/000-confi
|
||||
| [[skills/]] | Skill definitions |
|
||||
| [[modes/]] | Model mode presets (creative, research, code, etc.) — incl. `workplace-modes.md` |
|
||||
| [[workspace/]] | Workspace layouts and setup |
|
||||
| [[notebooks/]] | Jupyter / Colab notebooks (runnable .ipynb artifacts) |
|
||||
| [[settings/]] | Open WebUI / server settings |
|
||||
|
||||
## Migration Note
|
||||
|
||||
@@ -0,0 +1,32 @@
|
||||
---
|
||||
title: "Notebooks"
|
||||
status: "active"
|
||||
folder: "000-configs/notebooks"
|
||||
tags: [readme, notebook, code]
|
||||
created: "2026-07-30"
|
||||
updated: "2026-07-30"
|
||||
version: "1.0.0"
|
||||
---
|
||||
|
||||
# Notebooks
|
||||
|
||||
Executable Jupyter / Colab notebooks that accompany research projects or workspace workflows. These are *runnable* artifacts (.ipynb), not docs — the companion writeup lives in `06-research/` or as a task in `07-tasks/`.
|
||||
|
||||
## How to use
|
||||
|
||||
1. Download the `.ipynb` from this folder (or `git pull` on your PC and open in Obsidian / Jupyter / Colab).
|
||||
2. Open in **Google Colab** for the Cloud-Drive-backed workflows (most notebooks here assume Colab free tier + Drive mount).
|
||||
3. Update the task file in `07-tasks/` when you re-run or modify.
|
||||
|
||||
## Current notebooks
|
||||
|
||||
| Notebook | Purpose | Companion task |
|
||||
|---|---|---|
|
||||
| [[docx_to_markdown_openwebui.ipynb]] | DOCX + MD → clean Markdown for Open WebUI knowledge bases (mammoth + python-docx fallback + fuzzy dedup + URL extraction) | [[07-tasks/docx-to-markdown-converter-notebook]] |
|
||||
|
||||
## When adding a new notebook
|
||||
|
||||
1. Drop the `.ipynb` here.
|
||||
2. Add a row to the table above with a link.
|
||||
3. Create or update a task in `07-tasks/<slug>.md` describing what it does and how to run it.
|
||||
4. Commit with `notebook: <one-line>` per the repo commit-message convention.
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,102 @@
|
||||
---
|
||||
title: "Task: DOCX→Markdown converter notebook for Open WebUI"
|
||||
status: "active"
|
||||
folder: "07-tasks"
|
||||
tags: [task, notebook, openwebui, docx, markdown, converter]
|
||||
created: "2026-07-30"
|
||||
updated: "2026-07-30"
|
||||
version: "1.0.0"
|
||||
priority: "high"
|
||||
due: "2026-07-30"
|
||||
owner: "stateofshit"
|
||||
---
|
||||
|
||||
# Task: DOCX→Markdown converter notebook for Open WebUI
|
||||
|
||||
## Done — 2026-07-30
|
||||
|
||||
## What
|
||||
|
||||
A Colab notebook that batches `.docx` (and `.md`) inputs into clean Markdown optimized for embedding in **Open WebUI** knowledge bases. Open WebUI's RAG uses Markdown Header Splitting, so the converter preserves a clean single-H1 / nested-H2-H3 heading hierarchy — that's what makes retrieval chunks good. Output is mirrored to Google Drive so nothing is lost across Colab wipes.
|
||||
|
||||
This notebook is the upgrade over the older `/home/user/00-incoming/docx_to_markdown.ipynb`: it adds the Gemini-Notebook setup nuance (cite markers stripped, hard bold-promoted to headings) and the dedup edge-cases we hit this week.
|
||||
|
||||
## Pipeline (per file)
|
||||
|
||||
- **DOCX**: `mammoth(html) → clean_markdown → strip_masthead → strip_cite_markers → collapse_hr → promote_structure → collapse_extra_h1 → conditional synthetic H1 prepend`
|
||||
- **MD** (raw uploaded): same chain minus mammoth
|
||||
- **Fallback**: if mammoth crashes (the `_accept0` bug on weird docx), the pipeline degrades to a `python-docx` parser that still produces clean markdown
|
||||
- **Output**: one `.md` per input → `/drive/MyDrive/to_convert/markdown_output/`
|
||||
|
||||
## Headlining features
|
||||
|
||||
| # | Feature | Why it matters |
|
||||
|---|---|---|
|
||||
| 1 | Drive-backed manifest dedup (SHA-256 of input bytes) | Renaming a file does NOT cause a reconvert. Survives Colab wipes. |
|
||||
| 2 | **Pre-flight fuzzy dedup** (new today — cell 5b) | Drops `(1)`/`(2)`/`Copy of` variants where the bytes differ by a few chars but the doc is really the same. Uses Jaccard similarity over 500-char normalized text chunks. Threshold `0.97` (tunable). |
|
||||
| 3 | Gemini masthead + cite-marker stripping | Removes `— / Created By: Due Process Demolition / APP: Gemini Notebook / DATE: …` boilerplate and `[cite: N]` artifacts so the KB doesn't fill with garbage |
|
||||
| 4 | Auto heading promotion | For Gemini-Notebook-style docs that use bold paragraphs (not Word heading styles) for section titles, the converter promotes `**Bold**` lines to real `##`/`###` so Open WebUI's splitter has structure |
|
||||
| 5 | Single-H1 enforcement | `collapse_extra_h1()` keeps the first H1 (title) and demotes any subsequent H1 → H2. Required by Open WebUI's Markdown Header Splitter |
|
||||
| 6 | URL extraction | Every `http(s)://` URL found in converted .md files → `sources.csv` paired with source-doc SHA + stem |
|
||||
| 7 | Bulk downloader (off by default, cell 11) | Once sources.csv is curated, opt-in fetch of PDFs first then HTML to `/downloads/`, tracked in `downloads.csv` |
|
||||
| 8 | Zip + download (cell 10) | Bundle all converted `.md` as a single zip for upload to Open WebUI |
|
||||
|
||||
## File
|
||||
|
||||
| File | Purpose |
|
||||
|---|---|
|
||||
| `000-configs/notebooks/docx_to_markdown_openwebui.ipynb` | The notebook — 29 cells, nbformat-valid |
|
||||
| `000-configs/notebooks/README.md` | Index of all notebooks, how to add new ones |
|
||||
|
||||
## How to run (final user workflow)
|
||||
|
||||
1. Open the `.ipynb` in **Colab free tier** → `Runtime → Run all` (or run cell-by-cell).
|
||||
2. Cell 2: drop the docx/md files (Colab upload picker), or it auto-extracts from `/drive/MyDrive/zzz-new_shit/*.tar`.
|
||||
3. Cell 5b (idempotent): watch the pre-flight dedup report — see which `(1)`/`(2)` dupes got dropped by content overlap.
|
||||
4. Cell 6: conversions run; manifest is updated with statuses (`converted` / `crashed_fallback_ok` / `crashed_skipped`).
|
||||
5. Cell 7: `sources.csv` gets all URLs from the new conversions.
|
||||
6. Cell 8: heading-structure audit (counts H1/H2/H3 per file — sanity check for Open WebUI chunking).
|
||||
7. Cell 10: zip + download everything → upload zip into Open WebUI knowledge base.
|
||||
|
||||
## Drive layout (created on first run)
|
||||
|
||||
```
|
||||
/drive/MyDrive/to_convert/
|
||||
├── manifest.csv ← SHA-256 dedup history (append-only)
|
||||
├── dedup_drops.csv ← pre-flight fuzzy-dedup log (new today)
|
||||
├── archive/ ← original input bytes, durable across Colab wipes
|
||||
├── markdown_output/ ← converted .md files (this is what you upload)
|
||||
├── sources.csv ← doc → URLs (append-only)
|
||||
└── downloads/
|
||||
├── pdf/ html/ other/
|
||||
└── downloads.csv ← URL → saved file (append-only)
|
||||
```
|
||||
|
||||
## Config (cell 9 flags — tunable without touching the converter)
|
||||
|
||||
```python
|
||||
SKIP_ALREADY_SEEN = True
|
||||
ARCHIVE_INPUTS = True
|
||||
DEDUP_SIMILARITY_THRESHOLD = 0.97 # set to 0 to disable fuzzy dedup
|
||||
DEDUP_CHUNK_SIZE = 500 # chars per chunk in the similarity signature
|
||||
EXTRACT_IMAGES = True # Extract media to per-doc media/ (no embedding signal yet)
|
||||
```
|
||||
|
||||
## Bug fixed this session
|
||||
|
||||
- **TypeError: 'NoneType' object is not subscriptable** in cell 6 when manifest has a hash in the seen-set but `is_seen()` returns None (race with `manifest.csv` written by another process). Guarded with `seen_row = is_seen(sha, manifest) or {}` and `.get(..., '(unknown)')` accessors.
|
||||
|
||||
## Outstanding
|
||||
|
||||
- [ ] Re-run the notebook end-to-end with the user's real upload batch (the 257-file set with the `clean_email_threads.docx` + `clean_email_threads (2).docx` 25-byte-diff pair); confirm the fuzzy dedup actually drops the `(2)` variant at threshold 0.97.
|
||||
- [ ] If threshold is too generous (drops genuinely different docs), bump to `0.95`. If too strict (lets dupes through), drop to `0.90`. The user has the knob in cell 9 — no code changes needed.
|
||||
- [ ] Confirm `manifest.csv`, `dedup_drops.csv`, `sources.csv` all appear on Drive after first run and round-trip correctly across Colab sessions.
|
||||
- [ ] Consider adding a MinHash variant if Jaccard over fixed chunks ever gets slow on >1k inputs.
|
||||
|
||||
## Build provenance
|
||||
|
||||
The notebook is built from a Python script (`/tmp/build_notebook.py`) that stitches fragments together and is then validated. That script is a build artifact and lives outside the vault (single-use); future revisions to the notebook can be made directly by editing the `.ipynb` in Colab and committing the saved file, OR by patching the build script and rerunning. The fragment files were in `/tmp/nbfrags/` at build time — restore from git history if rebuilding.
|
||||
|
||||
## Related
|
||||
|
||||
- [[06-research/gemini-notebook-poweruser-briefing]] — context for the Gemini-Notebook .docx format quirks (cite markers, masthead, bold-as-heading) that this converter handles
|
||||
Reference in New Issue
Block a user