Files
Art-gallery/Documentation/data-and-images.md
T
Danila KhodjaefandCursor f78c14f307 Fix debug upload persistence and UX; exclude prod audit log from DB restore
Uploads and fixes now bust browser cache via file-mtime keys in API
responses. Debug upload shows a centered loading overlay and blocks search
while uploading. Prod DB restore skips curator_audit_log.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-09 18:30:19 +03:00

507 lines
32 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Art Gallery — data and images
How catalog content, biographies, and artwork files enter the system.
## Principles
1. **No runtime hot-linking** — the UI reads from `/images/…` (local disk). External URLs are used only during ingest.
2. **No AI-generated art or text** — biographies and descriptions come from Wikipedia; influence notes from curated art-history sources.
3. **Local copies** — every displayed image should exist under `data/images/` after seeding or fetch.
## Directory layout
```text
data/images/
├── portraits/ # Artist headshots (display ~900px wide)
│ ├── Claude_Monet.jpg
│ └── thumbs/ # Timeline thumbnails (~256px)
│ └── Claude_Monet_thumb.jpg
└── paintings/
├── Claude_Monet_Water_Lilies.jpg
└── thumbs/
└── Claude_Monet_Water_Lilies_thumb.jpg
```
File names are sanitised `{Artist}_{Title}.{ext}`. The image service can rediscover files on disk even when DB paths are empty (`server/image-service.js``syncPaintingFromDisk`).
### Dev vs production image storage
| Environment | Path on disk | Sync |
|-------------|--------------|------|
| **Development** | `./data/images/` in repo | Working copy on dev PC |
| **Production** | `/mnt/BasePool/Applications/Gallery/data/images` on TrueNAS | SMB `\\192.168.10.122\Gallery\data\images` |
Promote dev → prod files: `npm run devtoprod:images` (after `net use \\192.168.10.122\Gallery`). Refresh dev from prod: `npm run prodto:dev:images`. See [environments.md](environments.md).
## Scripts overview
| Script | npm command | Role |
|--------|-------------|------|
| `seed-wikipedia.js` | `npm run dev:seed` | Initial eras, movements, artists, one flagship painting per artist |
| `seed-catalog-data.js` | *(data only)* | Eras, movements, artist metadata consumed by seed |
| `sync-image-paths.js` | `npm run dev:sync-image-paths` | Import painting rows from disk; set `image_path` / `thumbnail_path` |
| `fetch-artist-images.js` | `npm run dev:fetch-artist-images` | Download or link artist portraits under `data/images/portraits/` |
| `fetch-artist-bios.js` | `npm run dev:fetch-artist-bios` | Wikipedia intros → `bio_short` / `bio_full` |
| `famous-paintings-data.js` | *(data only)* | Curated list of notable works per artist |
| `expand-paintings.js` | `npm run dev:expand-catalog` | Inserts works from data file for thin catalogs |
| `art-influences-data.js` | *(data only)* | Curated influence edges (painting / artist / movement) |
| `update-influences.js` | `npm run dev:update-influences` | Applies influence graph; creates missing artists/works |
| `fetch-missing-images.js` | `npm run dev:fetch-images` | Downloads files for paintings missing on disk |
| `image-fetcher.js` | *(library)* | Wikimedia / museum resolution used by fetch scripts and API |
| `sync-images-to-prod.ps1` / `sync-images-from-prod.ps1` | `npm run devtoprod:images` / `npm run prodto:dev:images` | Robocopy via SMB `\\192.168.10.122\Gallery` |
| `regenerate-thumbnails.js` | `npm run dev:regenerate-thumbnails` | Rebuild painting thumbs from full images via `sharp` |
| `regenerate-portrait-thumbs.js` | `npm run dev:regenerate-portrait-thumbs` | Rebuild timeline portrait thumbs (~256px) and set `portrait_thumb_path` |
| `audit-painting-images.js` | `npm run dev:audit-painting-images` | Detect thumb/full aspect-ratio mismatches |
| `find-duplicate-paintings.js` | `npm run dev:find-duplicates` | Report exact and near-duplicate catalog rows |
| `migrate-checkup-flags.js` | `npm run dev:migrate:checkup-flags` | Add `checkup_checked` / `checkup_fixed` columns |
## Typical workflow
```text
migrate → seed → sync-image-paths → fetch-artist-images → fetch-artist-bios → expand-catalog → update-influences → fetch-images (per artist or batch) → build client
```
1. **Seed** creates the base catalog (one flagship painting per artist; ~100 artists).
2. **sync-image-paths** imports additional paintings when `data/images/paintings/` already contains files from a full clone (filename pattern `{Artist}_{Title}.jpg`).
3. **fetch-artist-images** sets `portrait_path` from local files or Wikipedia.
4. **fetch-artist-bios** fills biography fields for every artist with a `wikipedia_title`.
5. **expand-catalog** brings each artist up to at least **6** notable works (configurable via `MIN_PAINTINGS`).
6. **update-influences** loads the influence graph (*Influenced By* / *Influenced* panels, 3D hall lamps, exit navigation).
7. **fetch-images** downloads artwork files still missing on disk; the 3D gallery needs local files for reliable textures.
## Seeding pipeline
`npm run dev:seed` runs `scripts/seed-wikipedia.js`, which:
1. Inserts **historical eras** and **art movements** (curated date ranges and colours from `seed-catalog-data.js`).
2. For each curated **artist**:
- Creates **artist periods** and one **flagship painting**.
- May download portraits and painting images when run with `--fetch-images`.
3. Does **not** insert influence edges — run `npm run dev:update-influences` after seed (see [Painting influence graph](#painting-influence-graph)).
Those influence edges power **3D hall navigation** and painting detail panels via **`painting_influence_sources`** (see `GET /api/artists/:id/navigation` and `GET /api/paintings/:id` in [API.md](API.md)).
Artists are grouped by movement and century; the seed list targets at most ~100 artists per century.
## Artist biographies
`npm run dev:fetch-artist-bios` reads each artists `wikipedia_title`, fetches the English Wikipedia **lead section**, and stores:
| Field | Content |
|-------|---------|
| `bio_short` | First two sentences |
| `bio_full` | Full intro (text before the first section heading) |
Flags:
- `--force` — refresh bios even when `bio_full` is already set.
**Disambiguation and title overrides** live in `ARTIST_WIKI_OVERRIDES` inside `scripts/fetch-artist-bios.js`:
| Artist in DB | Wikipedia article used |
|--------------|------------------------|
| Zeuxis | `Zeuxis (painter)` |
| Ivan Klyun | `Ivan Kliun` |
| Jean-Antoine Watteau | `Antoine Watteau` |
The script detects disambiguation pages (“X may refer to:”) and tries fallbacks such as `{name} (painter)` before giving up. Requests are throttled (~3.5 s apart) with retry on HTTP 429.
The bio page (`ArtistBio.tsx`) shows lifespan, movement, summary, full text, and the source Wikipedia title.
## Expanding thin catalogs
Many seed artists arrive with only one famous painting. `npm run dev:expand-catalog` runs `scripts/expand-paintings.js`, which:
1. Finds artists with fewer than `MIN_PAINTINGS` (default **6**).
2. Inserts rows from `scripts/famous-paintings-data.js` that are not already present (normalized title matching skips duplicates).
3. Sets `wikipedia_title` on each new painting for image resolution.
```bash
npm run dev:expand-catalog # DB rows only
npm run dev:expand-catalog -- --fetch-images # also download images (very slow)
```
To add more works, append entries to `famous-paintings-data.js`:
```javascript
{ artist: 'Gustav Klimt', title: 'Portrait of Adele Bloch-Bauer I', year: 1907 },
{ artist: 'Gustav Klimt', title: 'The Kiss', year: 1908, wikipedia_title: 'The Kiss (Klimt painting)' },
```
`wikipedia_title` is optional; it defaults to `title`. Use it when the Wikipedia article name differs from the display title.
Renaissance and medieval masters with large museum catalog dumps (e.g. Raphael, Dürer) are usually above the minimum already when **`sync-image-paths`** has imported files from disk; expansion targets Impressionists, modernists, and other artists who had only a single seed painting.
## Importing paintings from disk
When the repository includes a full `data/images/paintings/` tree but the database was seeded fresh (one row per artist), run:
```bash
npm run dev:sync-image-paths
```
`scripts/sync-image-paths.js`:
1. Scans `data/images/paintings/` for full-size files (not `thumbs/`).
2. Matches filenames to artists using the same `{Artist}_{Title}` sanitisation as `server/image-service.js`.
3. **Updates** `image_path` / `thumbnail_path` on existing rows when files are found.
4. **Inserts** missing painting rows for files not yet in the catalog.
Flags:
- `--dry-run` — report counts only, no DB writes.
Safe to re-run; already-imported works are skipped by normalized title matching.
Typical result on a full clone: ~1,000+ paintings linked from ~1,000 on-disk files.
## Artist portraits
Timeline movement flow loads **`portrait_thumb_path`** (~256px JPEG under `portraits/thumbs/{Artist}_thumb.jpg`) when available; biography and 3D exit navigation use full `portrait_path`. After adding portraits, run `npm run dev:regenerate-portrait-thumbs` to backfill thumbs on dev.
`npm run dev:fetch-artist-images` runs `scripts/fetch-artist-images.js`:
1. For each artist, checks `data/images/portraits/{Artist}.jpg` (or other extensions) and sets `portrait_path` when a local file exists.
2. Otherwise downloads from Wikipedia / search fallbacks via `image-fetcher.js`.
Flags:
- `--force` — re-fetch even when `portrait_path` is already set.
- `--limit=N` — process only the first N artists needing portraits.
Run after seed when portrait files exist on disk but the DB still has null `portrait_path` values.
## Painting influence graph
Directed influence links are stored in **`painting_influence_sources`**. Each row connects a painting to a **source** of type `painting`, `artist`, or `movement`, with optional period context (e.g. influence during the works creation year).
The legacy **`painting_influences`** table (painting-to-painting only) is still written alongside sources when running `npm run dev:update-influences` — it keeps script compatibility and matches the backfill migration. **The API reads only `painting_influence_sources`**, so each edge appears once in the UI.
Influence data drives:
- **Painting detail** — *Influenced By* (left) and *Influenced* (right) panels: painting thumbnails, artist portraits, or movement colour swatches, plus notes, aspects, period labels, and citations
- **3D hall exit** — predecessor and successor artists grouped by movement (from painting and artist sources)
- **3D gallery lamps** — golden picture light above frames with any influence edge (`has_influence_links` on painting API responses)
### Audit duplicate influence links
Painting-to-painting edges exist in both tables by design. To confirm the database has no stray duplicates and that the API model is clean:
```bash
npm run dev:audit-influence-duplicates
```
Reports: edges present in both tables, duplicate rows within either table (should be 0), and legacy-only / sources-only mismatches. If *Influenced* ever shows the same successor twice, restart the server after pulling API fixes — responses must not union legacy and sources tables.
### One-time migration
```bash
npm run dev:migrate:influence-sources # create table + backfill legacy painting edges
```
### Curated updates
`npm run dev:update-influences` runs `scripts/update-influences.js` against `scripts/art-influences-data.js`:
```bash
npm run dev:update-influences # insert curated edges
npm run dev:update-influences -- --fetch-images # also download images for newly created works
npm run dev:update-influences -- --discover # curated + web discovery pass
npm run dev:discover-influences # discovery only (no curated file pass)
npm run dev:update-influences -- --discover --limit=20 # cap discovery to N works
```
Each entry defines a later `work` and one or more `influencedBy` sources. Legacy single-object form is still supported:
```javascript
{ work: { artist: '...', title: '...', year: 1907 },
influencedBy: { artist: 'Giotto', title: 'Lamentation' } }
```
Multi-source form (painting, artist, movement):
```javascript
{
work: { artist: 'Pablo Picasso', title: "Les Demoiselles d'Avignon", year: 1907 },
influencedBy: [
{ type: 'painting', artist: 'Paul Cézanne', title: 'The Bathers' },
{ type: 'artist', artist: 'Paul Cézanne', period: { duringCreation: true } },
{ type: 'movement', movement: 'Fauvism', period: { start: 1905, end: 1907, note: '...' } },
],
notes, aspects, source_author, source, source_url,
}
```
When `artistMeta` is included, missing artists are created with movement and lifespan. Missing paintings are inserted with `wikipedia_title` for image fetch.
### Web discovery
`scripts/influence-discovery.js` searches art-history sources when `--discover` or `--discover-only` is passed:
- Wikipedia summaries and Wikidata **P737** (influenced by)
- Met Museum collection API
- DuckDuckGo site-restricted search across TheArtStory, Met, Google Arts & Culture, NGA, Art Institute of Chicago, MoMA, Britannica, JSTOR, Oxford Art Online, WikiArt, and Wikipedia
Discovered rows are stored with `confidence: discovered` and `discovered_via` (e.g. `wikipedia`, `wikidata`, `met`, `web:theartstory.org`). Period hints are inferred when the source text mentions influence during creation or a date range overlapping the works year.
Extend `art-influences-data.js` for high-quality curated chains; use discovery to suggest additional artist and movement links for manual review.
## Movement lineage (frontend flow diagram)
Separate from the painting influence graph, `client/src/data/movement-lineage.ts` lists **art-movement** predecessor→successor pairs used only by `MovementBands.tsx` to position streams and draw branch connectors (e.g. Impressionism → Post-Impressionism → Fauvism / Cubism / Expressionism).
| Aspect | Detail |
|--------|--------|
| Storage | TypeScript module in the client — **not** a database table |
| Format | `[parentMovementName, childMovementName]` tuples; names must match `art_movements.name` from the seed |
| Multiple parents | Allowed (e.g. Post-Impressionism feeding several modern paths) |
| Sources | Curator notes in the file reference Met essays, TheArtStory, and similar |
To add or fix a branch, edit `MOVEMENT_LINEAGE` in that file and rebuild the client. No migration or API change is required.
## Movement gallery interiors (frontend)
Separate from movement lineage layout, `client/src/data/movement-interior-styles.ts` defines a **unique 3D interior** for each seeded art movement (26 styles): wall/floor/ceiling textures, trim colours, window style, and architectural details (columns, coffered ceilings, etc.). Textures are generated procedurally in `client/src/utils/galleryProceduralTextures.ts`.
Wing layout (up to 55 works per wing, side-wall-only hang, window gap placement) lives in `client/src/utils/movementHallLayout.ts`. To change a movements look, edit its entry in `movement-interior-styles.ts` and rebuild the client.
## Historical event markers (frontend timeline)
`client/src/data/historical-events.ts` lists **world-history** markers shown on `Timeline.tsx` (French Revolution, World War I/II, etc.).
| Aspect | Detail |
|--------|--------|
| Storage | TypeScript module in the client — **not** a database table |
| Format | `{ id, name, startYear, endYear?, shortLabel? }` — omit `endYear` for a single-year pin |
| Interaction | Click a marker to zoom the shared timeline/movement view to that period |
| Year axis labels | Dynamic density in `Timeline.tsx` via `chooseTimelineTickInterval()` — fewer labels when zoomed out |
| Pan/zoom batching | `createViewChangeScheduler()` in `timelineView.ts` — one React update per animation frame |
| Vertical guides | `TimelineEventGuides.tsx` draws faint gold lines (or shaded spans) from the marker row down through the movement flow, aligned to the same year scale |
Edit `HISTORICAL_EVENTS` and rebuild the client to extend the set.
## Painting annotations (art-history notes)
Short curator-style notes on the painting detail page — separate from the influence graph.
| Aspect | Detail |
|--------|--------|
| Storage | PostgreSQL table `painting_annotations` |
| UI | `PaintingAnnotations.tsx` — numbered markers on the image (when `pos_x` / `pos_y` set) plus an “Art history notes” list |
| API | Included as `annotations[]` on `GET /api/paintings/:id` |
| Curated data | `scripts/painting-annotations-data.js` — artist/title keys matched via `influence-resolver.js` |
| Load | `npm run dev:update-painting-annotations` (replaces existing rows per painting by default) |
| Wikipedia pass | `npm run dev:update-painting-annotations -- --wikipedia` — one intro sentence per work from `wikipedia_title`; use `--wiki-delay=3000` if rate-limited; `--no-replace` to append without clearing curated rows |
Categories include `subject`, `technique`, `context`, and `symbolism`. Sources cite Gombrich, museum catalogs, and Wikipedia as appropriate.
## Batch image fetch
`npm run dev:fetch-images` (alias: `npm run dev:search-missing-paintings`) runs `scripts/fetch-missing-images.js`. It searches multiple sources for paintings without local files:
| Source | Notes |
|--------|--------|
| Wikidata / Wikipedia | Article image + P18 property |
| Wikimedia Commons | Direct file + search |
| Wikipedia search | Discovers better article title when seed `wikipedia_title` is a museum catalog label |
| Google Arts & Culture | Search + asset pages (`artsandculture.google.com`) |
| Met Museum | Open Access API |
| Art Institute of Chicago | IIIF open access |
| Cleveland Museum of Art | CC0 API |
| Rijksmuseum | Public API |
| Musée du Louvre | Collections search (`collections.louvre.fr`) |
| Europeana | European museum aggregator (optional `EUROPEANA_API_KEY` in `.env`) |
| Smithsonian | Optional (`SMITHSONIAN_API_KEY` in `.env`) |
| Harvard Art Museums | Optional (`HARVARD_ART_API_KEY` in `.env`) |
```bash
npm run dev:fetch-images # all missing, catalog order (~hours)
npm run dev:fetch-images -- --limit=50 # random sample of 50; 10s max per painting
npm run dev:fetch-images -- --limit=250 # random sample up to N (caps at current missing count)
npm run dev:fetch-images -- --limit=50 --max-wait=120 # slower, more thorough lookup per painting
npm run dev:fetch-images -- --artist="Albrecht Dürer" # one artist, catalog order
npm run dev:fetch-images -- --discover-only --limit=20 # fix wikipedia_title only
npm run dev:fetch-images -- --web-search-only --limit=50 # DuckDuckGo + Commons + multilingual Wikipedia
```
Each run prints **`Missing local files: N`** at startup — that is the current count of catalogued paintings with no full-size or thumbnail file under `data/images/`. A painting counts as present if **either** file exists on disk.
| Flag | Effect |
|------|--------|
| `--limit=N` | Process at most **N** paintings. Queue is a **random sample** of all works missing local files (not alphabetical). |
| `--max-wait=N` | Stop each painting after **N** seconds (default **10**). Logs `⏱ timeout` and continues. |
| `--artist="Name"` | Only that artists missing works, in catalog order (`sort_order`, `year`). |
| `--discover-only` | Update `wikipedia_title` via search; no download. |
| `--web-search-only` | Skip museum APIs; use web search + Commons + multilingual Wikipedia. |
Each limited batch run stops per painting after **`--max-wait` seconds** (default **10**), including source lookup and download. While a batch deadline is active, inter-request throttling is skipped and each HTTP call times out at the **remaining** budget (not the full 15s on-demand limit). Override the default via `FETCH_MAX_WAIT_SEC` in `.env` or `--max-wait=N`. On-demand fetches in the web UI keep their separate 15s API timeout and are unaffected.
Re-run the same command to pick a new random batch until the missing count reaches zero.
## On-demand image resolution
When a painting has no local file, `GET /api/paintings/:id/image` triggers `ensurePaintingImages()`:
1. Check DB paths → verify file on disk.
2. Scan disk by `{artist}_{title}` pattern.
3. If still missing and `wikipedia_title` is set, call `scripts/image-fetcher.js`:
- Resolve overrides and simplified titles
- **Wikipedia search** when catalog labels fail
- Wikidata → Wikimedia Commons → Wikipedia page image
- Fallbacks: Met Museum, Art Institute of Chicago, Cleveland Museum, Rijksmuseum, Smithsonian*, Harvard*
4. Save full image, **generate thumbnail by resizing the full file** (not a separate Commons thumb URL), update DB, serve file.
Separate Wikipedia/Commons thumbnail URLs often resolve to the **wrong work** (e.g. a different painting with a similar title). Thumbnails are always derived locally from the downloaded full image via `sharp` in `scripts/image-fetcher.js`.
Requests are deduplicated (`inflight` map) and timeout after 15 seconds. On-demand resolution uses a ~2.5 s delay between external requests to reduce rate-limit risk; batch `fetch-images` runs skip that delay while the per-painting deadline is active.
## Preload before 3D gallery
`POST /api/artists/:id/preload-images` is a **public** route (no curator login). It runs **local-only** linking — no network. The React client calls it automatically when entering an **artists** 3D hall so textures use files already on disk.
**Movement galleries** (`GET /api/movements/:id/gallery`) do not use preload — they load the full painting list from the API and resolve local paths the same way as artist halls. Works without files still show the canvas cover in the frame.
The 3D scene uses `galleryImageUrl()`, which never hits the on-demand API (remote latency breaks WebGL texture loading).
## Placeholders
When no image is available:
- `/placeholder-portrait.svg` — timeline and movement-flow portraits
- `/placeholder-art.svg` — paintings in lists and detail view
- **3D gallery** — draped **canvas cover** inside the frame (`CanvasCover` in `VirtualGallery.tsx`); shown when there is no local file, the fetch failed, or the texture has not loaded yet
These live in `client/public/` (and `client/dist/` after build).
## Adding new artists manually
1. Insert rows into `artists`, `artist_periods`, `paintings` (or extend the seed script).
2. Place image files under `data/images/` using the naming convention.
3. Run `npm run dev:fetch-artist-bios` for the new artists biography.
4. Add entries to `famous-paintings-data.js` and run `npm run dev:expand-catalog` if needed.
5. Run `npm run dev:fetch-images -- --artist="…"` or rely on preload / on-demand sync.
6. Add influence rows via `npm run dev:update-influences` / `art-influences-data.js` (writes both `painting_influence_sources` and legacy painting edges).
## PainterPalette external dataset
`Inputs/PainterPalette.csv` is a curated dataset (~10,000 painters) merging WikiArt, Art500k, and Wikidata with cleaned influence fields. The gallery integrates it **for existing DB artists only** (not a full catalog import).
### One-time setup
```bash
npm run dev:migrate:artist-palette # adds artists.palette_metadata JSONB
npm run dev:import-painter-palette # enrich + influence links
```
### What gets imported
| PainterPalette column | Gallery use |
|-----------------------|-------------|
| Nationality, gender, styles, birth/death places, occupations, … | Stored in `artists.palette_metadata` |
| `birth_year`, `death_year` | Fills missing DB years only |
| `Influencedby`, `Teachers` | **Artist** or **movement** sources on that artist's paintings |
| `Influencedon`, `Pupils` | Reverse **artist** sources on the pupil/successor's paintings |
Influence rows are written to `painting_influence_sources` with `discovered_via = painter-palette` and do not duplicate curated Gombrich edges (unique index + `ON CONFLICT DO NOTHING`).
Name matching uses normalized strings plus aliases in `scripts/painter-palette-lib.js` (e.g. `Bronzino``Agnolo Bronzino`, `J. M. W. Turner``J.M.W. Turner`). Museum names, WikiArt tags, and dimension strings are filtered out.
### Commands
```bash
npm run dev:analyze-painter-palette # match report
npm run dev:import-painter-palette -- --dry-run
npm run dev:import-painter-palette -- --metadata-only
npm run dev:import-painter-palette -- --influences-only
```
Re-run `import-painter-palette` after adding gallery artists or updating the CSV; existing palette influence rows are skipped if already present.
## Catalog export
Export the full painting catalog as CSV:
```bash
npm run dev:export-paintings
```
Writes **`Output/paintings.csv`** with columns `artist`, `painting`, `year` (sorted by artist, year, title). The `Output/` folder is git-ignored by convention; regenerate after catalog changes.
## Image fetcher overrides
`scripts/image-fetcher.js` includes hand-maintained overrides for ambiguous Wikipedia titles and direct URLs (e.g. works whose Commons name does not match the article title, or when museum search returns the wrong work). Extend these maps when automated resolution fails:
| Map | Use when |
|-----|----------|
| `PAINTING_WIKI_OVERRIDES` | DB / seed title should resolve to a different Wikipedia or Wikidata label |
| `DIRECT_IMAGE_OVERRIDES` | You know the exact Commons URL (bypasses Met / Art Institute false matches) |
Examples already in the repo:
- `Self-Portrait Hesitating` → Kauffman, Wikimedia Commons (National Trust)
- `Cherubs of the Sistine Madonna` → Raphaels putti detail, Wikimedia Commons
- `Madonna and Child (Madonna della Seggiola)` / `Madonna della seggiola` → Raphaels tondo, Palazzo Pitti
- `Madonna and Child` (Raphael) → *Small Cowper Madonna*, National Gallery of Art
- `Job Cigarette Papers` → Mucha poster disambiguation
- `Charing Cross Bridge` → Derain (not Monet)
After adding an override, delete any wrong cached file under `data/images/paintings/` and re-run fetch or call the on-demand image endpoint for that painting.
## Duplicate paintings
The catalog can contain the same work more than once — usually from a **double import** (identical artist + title + year + image) or **Wikipedia scrape variants** (different article titles for one icon, e.g. Andrei Rublevs Trinity).
### Find duplicates
```bash
npm run dev:find-duplicates
```
Runs `scripts/find-duplicate-paintings.js`, which reports:
| Report | Rule |
|--------|------|
| **Exact duplicates** | Same `artist_id`, `title`, and `year` |
| **Normalized title duplicates** | Same artist, titles differing only in punctuation/spacing |
| **Rublev / Trinity cluster** | Known multi-entry example for manual merge |
As of a recent audit (~1200 paintings): **52 exact duplicate pairs** (52 removable rows), concentrated in **Hieronymus Bosch** (25), **Albrecht Dürer** (14), and **Domenico Ghirlandaio** (13). Duplicate copies typically share the same image file and have **no influence links**, so the higher id in each pair is safe to delete after review.
**Near-duplicates** (different titles, same work) need curator judgment — e.g. Rublev ids 36, 316, 320, 322, 323 all describe the Trinity icon under different Wikipedia labels; keep id **36** (`Trinity`, wiki `Trinity (Andrei Rublev)`). For confirmed duplicates, debug **Remove entry** on painting detail is faster than manual SQL; it deletes files and the row in one step.
`expand-paintings.js` skips inserts when normalized titles match, but duplicates can still appear if seed and expansion use different title strings or if influence discovery creates works independently.
## Debug and checkup image fix
When **Debug mode** is on (home header) or from the **Checkup** page:
**Show more** (home header checkbox, `client/src/utils/debugMode.ts`) — when debug mode is on, automatically opens the **More** modal on each painting detail or artist bio page load (same as clicking **More**).
1. **Search**`GET /api/paintings/:id/debug-image-search` (or `…/debug-portrait-search` for artists) tries Google Custom Search (if `GOOGLE_CSE_API_KEY` + `GOOGLE_CSE_CX` are set in `.env`), Google Arts & Culture, Google Images scrape, then DuckDuckGo (`searchGoogleImagesFirst` / `searchArtistPortraitFirst` in `scripts/image-fetcher.js`).
2. **More**`GET …/debug-image-search/more` or `…/debug-portrait-search/more` returns up to 20 ranked candidates (`searchPaintingImagesMany` / `searchArtistPortraitMany`). The modal shows each thumbnail with **resolution** when the search API provides dimensions; otherwise the client probes via `GET /api/debug/image-proxy`.
3. **Fix**`POST …/fix-image` or `…/fix-portrait` downloads the chosen URL via `downloadImageForFix``replacePaintingImageFromUrl` / `replaceArtistPortraitFromUrl` in `server/image-service.js`. The server **always regenerates thumbnails from the saved full image** (`writePaintingThumb` / `writePortraitThumb` via `sharp` — not the search-result thumb URL), updates `thumbnail_path` / `portrait_thumb_path`, and sets `checkup_fixed` + `checkup_checked`.
4. **Clear**`POST …/clear-image` or `…/clear-portrait` deletes local file(s), nulls DB paths, sets both flags. Cleared slots stay empty in the UI (no placeholder; `checkup_fixed` prevents on-demand refetch for paintings).
5. **Upload**`POST …/upload-image` or `…/upload-portrait` accepts a base64-encoded file in JSON (Express body limit **20 MB**; decoded image max **15 MB**), validates with `sharp`, writes to the standard filename under `data/images/`, and regenerates the matching thumbnail the same way as **Fix it**. The file picker uses a native `<label>` + hidden `<input>` (`DebugUploadButton.tsx`) so the `change` event is reliable on Windows/Chromium.
6. **Remove entry** (painting detail only) — `DELETE /api/paintings/:id` via `deletePainting()` in `server/image-service.js`: deletes image files, removes the DB row (cascade on influence/annotation tables), refetches artist/movement gallery data, remounts the 3D hall, and navigates to the next or previous catalog work with no confirmation dialog.
### Debug panel (painting detail and artist bio)
With debug mode on, `PaintingDetail.tsx` and `ArtistBio.tsx` show a bottom-left panel with search preview and action buttons:
| Button | API (paintings / portraits) | Effect |
|--------|----------------------------|--------|
| **Checked** | `PATCH …/checkup-flags` `{ "checked": true }` | Marks reviewed; gold frame (paintings) or gold portrait border (artists) |
| **Fix it** | `POST …/fix-image` / `…/fix-portrait` | Saves top search result to disk, **regenerates thumb from full image**, sets both flags, refreshes detail + gallery / timeline |
| **More** | `GET …/debug-*-search/more` then fix endpoint | Modal with 20 clickable results (resolution label under each thumb); pick one to replace (thumb regenerated from downloaded full) |
| **Clear** | `POST …/clear-image` / `…/clear-portrait` | Removes file(s), empty frame in UI |
| **Upload** | `POST …/upload-image` / `…/upload-portrait` | Local file picker (`DebugUploadButton`) → save full image + **auto-generated thumb**; centered **Loading…** overlay; clears preview and blocks search/fix while uploading |
| **Remove entry** | `DELETE /api/paintings/:id` | **Paintings only** — permanent delete + gallery refresh + catalog navigation |
The client passes `searchUrl`, `source`, and `thumbUrl` from search results to improve download reliability. After a fix, clear, upload, or remove, `HomePage` updates the gallery session. Image URLs include `?v=<mtime_ms>` from API **`image_cache_key`** / **`thumbnail_cache_key`** (file `mtime` on disk) so replaced files reload after a full page refresh even when the relative path is unchanged. In-session counters still bump immediately after a mutation.
Checkup **Search visible** queues search for filtered rows only (3 concurrent); it does not search the full catalog on load.
**Curator login required** for all debug/checkup UI and mutating API routes. Guests can browse and enter 3D halls normally; preload remains public. See [API.md — Authentication](API.md#authentication).
See [API.md](API.md#developer-image-audit) and [basics.md](basics.md#developer-tools-image-audit).