# Scheer Enterprise Translation Tool A web tool connecting to the **DeepL API (Growth plan)** for text, document and PDF translation, plus token-optimized JSON translation, with per-domain glossaries for Scheer Group, Scheer IDS, Scheer PAS, and Scheer IMC. Supported languages everywhere in the tool: **German, German (Swiss), English (British), English (American), and French.** ## Why there's a small server involved DeepL's API blocks direct calls from browser JavaScript (CORS) and your key should never be embedded in public front-end code. `server.js` is a tiny zero-dependency Node proxy: it serves the UI and relays requests to DeepL using an API key configured **only on the server**. The key never touches the browser, localStorage, or any client-side code. ## Local setup (single user, on your own machine) 1. Requires Node.js 18+ (uses built-in `fetch`/`FormData`/`Blob` — no `npm install` needed). 2. Get your DeepL Growth API key from your DeepL account (developer/API settings). 3. Copy `.env.example` to `.env` in this folder and set `DEEPL_API_KEY=...`. (Alternatively, export it as a real environment variable: `DEEPL_API_KEY=xxx node server.js`.) 4. Run: ``` node server.js ``` 5. Open **http://localhost:3000**. The usage indicator in the top-right corner of the header loads automatically (click the ↻ next to it to refresh anytime) and confirms the server is connected to DeepL. ## Deploying for shared, multi-user access on a real domain To make this available to everyone in the org at a URL like `https://translate.yourcompany.com` instead of just `localhost`, you need three things: a server to run it on, a domain pointed at that server, and a way to keep it running. None of this requires code changes — `server.js` already listens on all network interfaces, not just `localhost`. 1. **Get a server.** Any internal VM, cloud instance, or on-prem box that can stay running and is reachable from wherever your users are (office network, VPN, or the public internet, depending on your needs). 2. **Point your domain at it.** Create a DNS A/AAAA record for `translate.yourcompany.com` pointing at that server's IP. 3. **Put a reverse proxy in front of it** (e.g. Nginx or Caddy) to handle HTTPS and forward requests to the Node app on port 3000. Caddy is the simplest option — a `Caddyfile` with just: ``` translate.yourcompany.com { reverse_proxy localhost:3000 } ``` gets you automatic HTTPS via Let's Encrypt with no extra config. 4. **Keep the Node process running** with a process manager or container instead of a terminal session — see the Docker option below, or use something like `pm2` or a `systemd` service if you'd rather run it bare. 5. **Configure `.env` and `domains.json` once, on that server** — every user hitting the domain shares the same DeepL key and the same four domain glossaries. That's the intended design: one shared quota, one shared set of glossaries per Scheer entity, managed centrally. ### Running it in Docker ```bash docker build -t scheer-translation-tool . # domains.json needs to exist before the bind mount below, otherwise Docker # will create it as a directory instead of a file: touch domains.json docker run -d --name scheer-translation \ -p 3000:3000 \ --env-file .env \ -v "$(pwd)/domains.json:/app/domains.json" \ --restart unless-stopped \ scheer-translation-tool ``` Then point your reverse proxy at `localhost:3000` on that host as described above. `--restart unless-stopped` makes it survive reboots. ### Important: this tool has no built-in login Anyone who can reach the domain can translate documents (consuming your shared DeepL quota) and edit any domain's glossary — there's no user authentication built in. Depending on your needs: - Restrict it to your internal network or VPN (don't expose it on the public internet), and/or - Add HTTP Basic Auth at the reverse proxy layer (a couple of lines in Nginx or Caddy), and/or - Put it behind whatever SSO/gateway your org already uses for internal tools. ## Features - **Text Translation** — free text in/out, placeholder protection (`{{var}}`, `%s`, `{id}` preserved verbatim), auto-detect or manual source language. - **Document Translation** — upload `.docx`/`.pptx`/`.xlsx`/`.pdf`/`.txt`/`.html`, tracks DeepL's upload → poll → download flow, gives you a translated file to download. - **PDF Translation** — same upload → poll → download flow, dedicated to PDFs, with the full language selection. - **JSON Translation** — paste or upload JSON, only extracts and translates user-facing string values (skips keys, numbers, booleans, nulls, URLs, HTML/code, template tags, and key-like identifiers such as `status_code_404`), batches everything into as few DeepL requests as possible, and reconstructs the exact tree. Shows a stats panel (values scanned/sent/skipped, character savings). - **Domain glossaries** — a domain bar at the top of every page (always visible, independent of which tab you're on) lets you pick one of the four Scheer entities: Scheer Group, Scheer IDS, Scheer PAS, Scheer IMC. Each domain stores **two fixed-direction glossaries**: German → English and English → German, each with its own list of exact term pairs, created via DeepL's glossary API. Click **⚙ Domain settings** to manage the currently selected domain's two term lists. Every translate action (text, document, PDF, JSON) automatically applies the matching direction's glossary and shows a note confirming it was applied. ## How glossaries are stored (and why everyone sees the same ones) Glossary terms are **not** stored in the browser (no localStorage, no per-user data). Saving a term list calls DeepL's glossary API to create the actual glossary on DeepL's servers, and the server writes the resulting glossary ID plus a local copy of the terms to **`domains.json`** — a JSON file that lives next to `server.js` on whichever machine is running the server. Every request from every browser goes through that one server, which reads the same `domains.json`, so anyone visiting the site sees and edits the same shared German→English and English→German lists per domain. The settings modal also re-fetches from the server each time it's opened, so you'll see another teammate's most recent edits rather than a stale local copy. For a small internal tool with a handful of admins editing glossaries occasionally, a flat JSON file is a reasonable, low-maintenance choice — there's no database to run or back up. If this ever needs to support many people editing glossaries concurrently, or an audit trail of who changed what, that would be the point to move `domains.json` into a real database. ## Known limitations - **JSON skip rules are heuristic** (regex-based). Always review the output panel before shipping translated JSON to production. - **Glossaries are recreated (not edited in place) on every save.** DeepL's v2 glossaries are immutable, so saving a domain's German→English or English→German list deletes the old one and creates a new one with the full term list. - Only these two directions are supported for glossaries. A glossary only applies to a translation when the request is going exactly German→English or English→German with an explicit (non-auto-detect) source language — the glossary API requires the source language to be set. ## Files - `server.js` — proxy server (no external dependencies), reads `DEEPL_API_KEY` from `.env`/environment - `index.html` — the single-page UI (styled with the Scheer brand palette: red, black, greys, whitespace) - `domains.json` — auto-created on first run; stores each domain's two glossary IDs/metadata (see storage section above) - `.env.example` — copy to `.env` and fill in your key - `package.json` — for `npm start` convenience (nothing to install) - `Dockerfile`, `.dockerignore` — for containerized deployment (see above)