Add random queue sampling and a 10s per-painting deadline for fetch-images batches, cap HTTP timeouts to the remaining budget, and update docs plus newly fetched artwork files. Co-authored-by: Cursor <cursoragent@cursor.com>
235 lines
13 KiB
Markdown
235 lines
13 KiB
Markdown
# Art Gallery — data and images
|
||
|
||
How catalog content, biographies, and artwork files enter the system.
|
||
|
||
## Principles
|
||
|
||
1. **No runtime hot-linking** — the UI reads from `/images/…` (local disk). External URLs are used only during ingest.
|
||
2. **No AI-generated art or text** — biographies and descriptions come from Wikipedia; influence notes from curated art-history sources.
|
||
3. **Local copies** — every displayed image should exist under `data/images/` after seeding or fetch.
|
||
|
||
## Directory layout
|
||
|
||
```text
|
||
data/images/
|
||
├── portraits/ # Artist headshots
|
||
│ └── Claude_Monet.jpg
|
||
└── paintings/
|
||
├── Claude_Monet_Water_Lilies.jpg
|
||
└── thumbs/
|
||
└── Claude_Monet_Water_Lilies_thumb.jpg
|
||
```
|
||
|
||
File names are sanitised `{Artist}_{Title}.{ext}`. The image service can rediscover files on disk even when DB paths are empty (`server/image-service.js` → `syncPaintingFromDisk`).
|
||
|
||
## Scripts overview
|
||
|
||
| Script | npm command | Role |
|
||
|--------|-------------|------|
|
||
| `seed-wikipedia.js` | `npm run seed` | Initial eras, movements, artists, paintings, influences |
|
||
| `fetch-artist-bios.js` | `npm run fetch-artist-bios` | Wikipedia intros → `bio_short` / `bio_full` |
|
||
| `famous-paintings-data.js` | *(data only)* | Curated list of notable works per artist |
|
||
| `expand-paintings.js` | `npm run expand-catalog` | Inserts works from data file for thin catalogs |
|
||
| `art-influences-data.js` | *(data only)* | Curated painting-to-painting influence edges |
|
||
| `update-influences.js` | `npm run update-influences` | Applies influence graph; creates missing artists/works |
|
||
| `fetch-missing-images.js` | `npm run fetch-images` | Downloads files for paintings missing on disk |
|
||
| `image-fetcher.js` | *(library)* | Wikimedia / museum resolution used by fetch scripts and API |
|
||
| `regenerate-thumbnails.js` | `npm run regenerate-thumbnails` | Rebuild thumbs from full images via `sharp` |
|
||
| `audit-painting-images.js` | `npm run audit-painting-images` | Detect thumb/full aspect-ratio mismatches |
|
||
|
||
## Typical workflow
|
||
|
||
```text
|
||
migrate → seed → fetch-artist-bios → expand-catalog → update-influences → fetch-images (per artist or batch) → build client
|
||
```
|
||
|
||
1. **Seed** creates the base catalog (often one flagship painting per modern artist).
|
||
2. **fetch-artist-bios** fills biography fields for every artist with a `wikipedia_title`.
|
||
3. **expand-catalog** brings each artist up to at least **6** notable works (configurable via `MIN_PAINTINGS`).
|
||
4. **fetch-images** downloads artwork files; the 3D gallery needs local files for reliable textures.
|
||
|
||
## Seeding pipeline
|
||
|
||
`npm run seed` runs `scripts/seed-wikipedia.js`, which:
|
||
|
||
1. Inserts **historical eras** and **art movements** (curated date ranges and colours).
|
||
2. For each curated **artist**:
|
||
- Creates **artist periods** and **paintings**.
|
||
- May download portraits and painting images (depending on seed script version).
|
||
3. Writes **painting_influences** edges from curated scholarship references.
|
||
|
||
Those influence edges power **3D hall navigation**: predecessors and successors at each artist’s exit doorway are computed from this table (see `GET /api/artists/:id/navigation` in [API.md](API.md)).
|
||
|
||
Artists are grouped by movement and century; the seed list targets at most ~100 artists per century.
|
||
|
||
## Artist biographies
|
||
|
||
`npm run fetch-artist-bios` reads each artist’s `wikipedia_title`, fetches the English Wikipedia **lead section**, and stores:
|
||
|
||
| Field | Content |
|
||
|-------|---------|
|
||
| `bio_short` | First two sentences |
|
||
| `bio_full` | Full intro (text before the first section heading) |
|
||
|
||
Flags:
|
||
|
||
- `--force` — refresh bios even when `bio_full` is already set.
|
||
|
||
**Disambiguation and title overrides** live in `ARTIST_WIKI_OVERRIDES` inside `scripts/fetch-artist-bios.js`:
|
||
|
||
| Artist in DB | Wikipedia article used |
|
||
|--------------|------------------------|
|
||
| Zeuxis | `Zeuxis (painter)` |
|
||
| Ivan Klyun | `Ivan Kliun` |
|
||
| Jean-Antoine Watteau | `Antoine Watteau` |
|
||
|
||
The script detects disambiguation pages (“X may refer to:”) and tries fallbacks such as `{name} (painter)` before giving up. Requests are throttled (~3.5 s apart) with retry on HTTP 429.
|
||
|
||
The bio page (`ArtistBio.tsx`) shows lifespan, movement, summary, full text, and the source Wikipedia title.
|
||
|
||
## Expanding thin catalogs
|
||
|
||
Many seed artists arrive with only one famous painting. `npm run expand-catalog` runs `scripts/expand-paintings.js`, which:
|
||
|
||
1. Finds artists with fewer than `MIN_PAINTINGS` (default **6**).
|
||
2. Inserts rows from `scripts/famous-paintings-data.js` that are not already present (normalized title matching skips duplicates).
|
||
3. Sets `wikipedia_title` on each new painting for image resolution.
|
||
|
||
```bash
|
||
npm run expand-catalog # DB rows only
|
||
npm run expand-catalog -- --fetch-images # also download images (very slow)
|
||
```
|
||
|
||
To add more works, append entries to `famous-paintings-data.js`:
|
||
|
||
```javascript
|
||
{ artist: 'Gustav Klimt', title: 'Portrait of Adele Bloch-Bauer I', year: 1907 },
|
||
{ artist: 'Gustav Klimt', title: 'The Kiss', year: 1908, wikipedia_title: 'The Kiss (Klimt painting)' },
|
||
```
|
||
|
||
`wikipedia_title` is optional; it defaults to `title`. Use it when the Wikipedia article name differs from the display title.
|
||
|
||
Renaissance and medieval masters with large museum catalog dumps (e.g. Raphael, Dürer) are usually above the minimum already; expansion targets Impressionists, modernists, and other artists who had only a single seed painting.
|
||
|
||
## Painting influence graph
|
||
|
||
Directed edges in `painting_influences` drive:
|
||
|
||
- **Painting detail** — *Influenced By* (left) and *Influenced* (right) panels with notes, aspects, and citations
|
||
- **3D hall exit** — predecessor and successor artists grouped by movement
|
||
- **3D gallery lamps** — a golden picture light above any frame whose work has an influence edge (`has_influence_links` on painting API responses)
|
||
|
||
`npm run update-influences` runs `scripts/update-influences.js` against `scripts/art-influences-data.js`:
|
||
|
||
```bash
|
||
npm run update-influences # insert edges; create missing artists/paintings
|
||
npm run update-influences -- --fetch-images # also download images for newly created works
|
||
```
|
||
|
||
Each entry defines a later `work` influenced by an earlier `influencedBy` painting, plus optional curator fields (`notes`, `aspects`, `source_author`, `source`). When `artistMeta` is included, missing artists are created with movement and lifespan. Missing paintings are inserted with `wikipedia_title` for image fetch.
|
||
|
||
Extend the data file to add lineage chains (e.g. Giotto → Masaccio → Michelangelo → Manet → Picasso → Warhol). New bridge artists may include Géricault, Friedrich, Constable, Giorgione, Poussin, de Chirico, and Böcklin.
|
||
|
||
## Batch image fetch
|
||
|
||
`npm run fetch-images` (alias: `npm run search-missing-paintings`) runs `scripts/fetch-missing-images.js`. It searches multiple sources for paintings without local files:
|
||
|
||
| Source | Notes |
|
||
|--------|--------|
|
||
| Wikidata / Wikipedia | Article image + P18 property |
|
||
| Wikimedia Commons | Direct file + search |
|
||
| Wikipedia search | Discovers better article title when seed `wikipedia_title` is a museum catalog label |
|
||
| Google Arts & Culture | Search + asset pages (`artsandculture.google.com`) |
|
||
| Met Museum | Open Access API |
|
||
| Art Institute of Chicago | IIIF open access |
|
||
| Cleveland Museum of Art | CC0 API |
|
||
| Rijksmuseum | Public API |
|
||
| Musée du Louvre | Collections search (`collections.louvre.fr`) |
|
||
| Europeana | European museum aggregator (optional `EUROPEANA_API_KEY` in `.env`) |
|
||
| Smithsonian | Optional (`SMITHSONIAN_API_KEY` in `.env`) |
|
||
| Harvard Art Museums | Optional (`HARVARD_ART_API_KEY` in `.env`) |
|
||
|
||
```bash
|
||
npm run fetch-images # all missing, catalog order (~hours)
|
||
npm run fetch-images -- --limit=50 # random sample of 50; 10s max per painting
|
||
npm run fetch-images -- --limit=920 # random sample up to N missing works
|
||
npm run fetch-images -- --limit=50 --max-wait=120 # slower, more thorough lookup per painting
|
||
npm run fetch-images -- --artist="Albrecht Dürer" # one artist, catalog order
|
||
npm run fetch-images -- --discover-only --limit=20 # fix wikipedia_title only
|
||
npm run fetch-images -- --web-search-only --limit=50 # DuckDuckGo + Commons + multilingual Wikipedia
|
||
```
|
||
|
||
| Flag | Effect |
|
||
|------|--------|
|
||
| `--limit=N` | Process at most **N** paintings. Queue is a **random sample** of all works missing local files (not alphabetical). |
|
||
| `--max-wait=N` | Stop each painting after **N** seconds (default **10**). Logs `⏱ timeout` and continues. |
|
||
| `--artist="Name"` | Only that artist’s missing works, in catalog order (`sort_order`, `year`). |
|
||
| `--discover-only` | Update `wikipedia_title` via search; no download. |
|
||
| `--web-search-only` | Skip museum APIs; use web search + Commons + multilingual Wikipedia. |
|
||
|
||
Each limited batch run stops per painting after **`--max-wait` seconds** (default **10**), including source lookup and download. While a batch deadline is active, inter-request throttling is skipped and each HTTP call times out at the **remaining** budget (not the full 15s on-demand limit). Override the default via `FETCH_MAX_WAIT_SEC` in `.env` or `--max-wait=N`. On-demand fetches in the web UI keep their separate 15s API timeout and are unaffected.
|
||
|
||
Re-run the same command to pick a new random batch until the missing count reaches zero.
|
||
|
||
## On-demand image resolution
|
||
|
||
When a painting has no local file, `GET /api/paintings/:id/image` triggers `ensurePaintingImages()`:
|
||
|
||
1. Check DB paths → verify file on disk.
|
||
2. Scan disk by `{artist}_{title}` pattern.
|
||
3. If still missing and `wikipedia_title` is set, call `scripts/image-fetcher.js`:
|
||
- Resolve overrides and simplified titles
|
||
- **Wikipedia search** when catalog labels fail
|
||
- Wikidata → Wikimedia Commons → Wikipedia page image
|
||
- Fallbacks: Met Museum, Art Institute of Chicago, Cleveland Museum, Rijksmuseum, Smithsonian*, Harvard*
|
||
4. Save full image, **generate thumbnail by resizing the full file** (not a separate Commons thumb URL), update DB, serve file.
|
||
|
||
Separate Wikipedia/Commons thumbnail URLs often resolve to the **wrong work** (e.g. a different painting with a similar title). Thumbnails are always derived locally from the downloaded full image via `sharp` in `scripts/image-fetcher.js`.
|
||
|
||
Requests are deduplicated (`inflight` map) and timeout after 15 seconds. On-demand resolution uses a ~2.5 s delay between external requests to reduce rate-limit risk; batch `fetch-images` runs skip that delay while the per-painting deadline is active.
|
||
|
||
## Preload before 3D gallery
|
||
|
||
`POST /api/artists/:id/preload-images` runs **local-only** linking — no network. Call this when entering an artist’s 3D hall so textures use files already on disk.
|
||
|
||
The 3D scene uses `galleryImageUrl()`, which never hits the on-demand API (remote latency breaks WebGL texture loading).
|
||
|
||
## Placeholders
|
||
|
||
When no image is available:
|
||
|
||
- `/placeholder-portrait.svg` — timeline / movement band portraits
|
||
- `/placeholder-art.svg` — paintings in lists and detail view
|
||
- **3D gallery** — draped **canvas cover** inside the frame (`CanvasCover` in `VirtualGallery.tsx`); shown when there is no local file, the fetch failed, or the texture has not loaded yet
|
||
|
||
These live in `client/public/` (and `client/dist/` after build).
|
||
|
||
## Adding new artists manually
|
||
|
||
1. Insert rows into `artists`, `artist_periods`, `paintings` (or extend the seed script).
|
||
2. Place image files under `data/images/` using the naming convention.
|
||
3. Run `npm run fetch-artist-bios` for the new artist’s biography.
|
||
4. Add entries to `famous-paintings-data.js` and run `npm run expand-catalog` if needed.
|
||
5. Run `npm run fetch-images -- --artist="…"` or rely on preload / on-demand sync.
|
||
6. Add influence rows to `painting_influences` with citation fields where possible.
|
||
|
||
## Image fetcher overrides
|
||
|
||
`scripts/image-fetcher.js` includes hand-maintained overrides for ambiguous Wikipedia titles and direct URLs (e.g. works whose Commons name does not match the article title, or when museum search returns the wrong work). Extend these maps when automated resolution fails:
|
||
|
||
| Map | Use when |
|
||
|-----|----------|
|
||
| `PAINTING_WIKI_OVERRIDES` | DB / seed title should resolve to a different Wikipedia or Wikidata label |
|
||
| `DIRECT_IMAGE_OVERRIDES` | You know the exact Commons URL (bypasses Met / Art Institute false matches) |
|
||
|
||
Examples already in the repo:
|
||
|
||
- `Self-Portrait Hesitating` → Kauffman, Wikimedia Commons (National Trust)
|
||
- `Cherubs of the Sistine Madonna` → Raphael’s putti detail, Wikimedia Commons
|
||
- `Madonna and Child (Madonna della Seggiola)` / `Madonna della seggiola` → Raphael’s tondo, Palazzo Pitti
|
||
- `Madonna and Child` (Raphael) → *Small Cowper Madonna*, National Gallery of Art
|
||
- `Job Cigarette Papers` → Mucha poster disambiguation
|
||
- `Charing Cross Bridge` → Derain (not Monet)
|
||
|
||
After adding an override, delete any wrong cached file under `data/images/paintings/` and re-run fetch or call the on-demand image endpoint for that painting.
|