Browser extension (Chrome, Edge, Firefox) that turns the captions a media page already carries into searchable, exportable text. It inspects a page only after the user clicks the toolbar action on it. When a page has no readable timed text, Chrome's built-in on-device model can transcribe the tab's audio or a dropped file. No backend, accounts, API keys, analytics, or paid services.
This file is the operating manual. Everything an operator or a future agent needs is here or linked from here.
| Store | Listing | Dashboard | API docs | Credentials (GitHub Actions secrets) | Notes |
|---|---|---|---|---|---|
| Chrome Web Store | ahddbfbjafmbceehebpeanpnlbaimepk | dashboard | docs | CWS_CLIENT_ID, CWS_CLIENT_SECRET, CWS_REFRESH_TOKEN, CWS_PUBLISHER_ID |
Listing text, screenshots and privacy answers are edited in the dashboard; answers are kept in store/privacy-fields.md |
| Microsoft Edge Add-ons | 069ca91d-a7cd-4bac-8224-1ee38a2d2a06 | dashboard | docs | EDGE_API_KEY, EDGE_CLIENT_ID |
API key expires 2026-11-29. Chromium build with manifest.edge.json: no tabCapture or offscreen, the built-in model has no audio input there. Renew the key at https://partner.microsoft.com/en-us/dashboard/microsoftedge/publishapi and update it in both extension repos |
| Firefox Add-ons (AMO) | video-transcript@qyl.at | dashboard | docs | AMO_JWT_ISSUER, AMO_JWT_SECRET |
One API key pair per Mozilla account, shared with the other extension repo; a new key invalidates the old one everywhere. The listing text is applied from store/listing.md |
The table is rendered from store.config.json by the
shared store-publish tool, which
save-media owns; this repository only consumes it and files issues there.
Change the config, then run bunx store-publish readme --write; CI fails when
the two drift. A new extension copies that file and changes the ids.
Credentials are GitHub Actions secrets in this repository. GitHub never returns their values; test those stored copies in an Actions run. The readable local copies listed below can also be used with the CLI. The same eight values are set on save-media, which shares the Chrome OAuth client, the Edge Publish API key, and the AMO key pair.
Two rules that have already cost a day each:
- Never generate a new AMO key. Mozilla issues one key pair per account. A new pair invalidates the old one immediately and breaks publishing in both repositories. An HTTP 401 almost always means the stored value is mangled, not that the key is gone.
- Never move a secret through the clipboard. A value read back with
pbpasteis not always what was copied. Check the length before writing: the Edge API key is 40 characters, the Edge client id and product id are 36.
Where the readable copies live on the owner's machine (presence only, never print values):
| Credential | Location |
|---|---|
| AMO key pair | macOS Keychain item AMO API (addons.mozilla.org); the issuer is the account field |
| Chrome OAuth client id and secret | ~/.config/vitals/cws-client.json |
| Chrome refresh token | ~/.config/vitals/cws-refresh-token.txt |
| Edge client id and API key | ~/.config/vitals/store-secrets.env, mode 600 |
| Register of all secret locations | ~/.config/vitals/keys.json |
| Script that writes the values into both repositories | ~/.config/vitals/set-store-secrets.sh |
Before using set-store-secrets.sh, populate all eight values from their
documented sources; the environment file alone does not supply the AMO pair.
For an Edge-only renewal, update only the two Edge secrets in each repository.
The Chrome OAuth client is "Desktop client 2" in Google Cloud project server
(uplifted-nuance-408417); its consent screen is in production.
gh workflow run store-status.yml -R ANcpLua/yt-transcript --ref main -f store=chrome
gh workflow run store-status.yml -R ANcpLua/yt-transcript --ref main -f store=firefox
gh workflow run store-status.yml -R ANcpLua/yt-transcript --ref main -f store=edge
gh workflow run store-status.yml -R ANcpLua/yt-transcript --ref main -f store=amo-listing-diff
gh run list -R ANcpLua/yt-transcript --workflow=store-status.yml --limit 1
gh run view <run id> -R ANcpLua/yt-transcript --logchromeprints the published revision and, while a review is open, the submitted revision with its state.firefoxlists every version with its review status.edgehas no read endpoint. The step probes the credentials with an operation id that cannot exist (HTTP 401 means rejected, anything else means accepted), prints the lengths of the three values, and prints the days left until the API key expires. It fails once the key has expired.amo-listing-diffcompares the## AMO descriptionsection ofstore/listing.mdwith the live AMO description.amo-listing-applywrites it. This is the only path that edits listing text; Chrome and Edge listing text is edited in their dashboards.
Versions live in manifest.json, manifest.edge.json,
manifest.firefox.json, and package.json, and must match. store-publish version refuses a mismatch
and, on a tag, a tag that differs from them.
One dispatch with stores=all attempts all three stores as sequential steps
in one job. This is not a parallel job matrix. Store acceptance and review
remain independent, so a successful submission is not proof that it is live.
# 1. bump all four version fields and add a CHANGELOG entry
bun install # refresh bun.lock after the version bump
bun run lint
bun run test
bun run build
bunx playwright test
bunx store-publish lint
bunx store-publish readme --check
release_version=$(bunx store-publish version)
# 2. stage the reviewed release changes, commit, push main, and wait for ci.yml
git add -p
git commit -m "Release $release_version"
git push origin main
# 3. after CI succeeds for this commit, create a new tag and push it
git tag "v$release_version"
git push origin "v$release_version"The tag run checks and builds the four zips and creates a GitHub release. It
does not submit to stores. Wait for that run to succeed. Before uploading,
run the status checks above and inspect Edge's dashboard for certification
state; its API probe checks credentials only. Resolve pending listing or
privacy fields and wait for existing reviews that block uploads. Confirm
main still contains the tested release code and version, then dispatch once:
gh workflow run release.yml -R ANcpLua/yt-transcript --ref main -f stores=all -f chrome=release
gh run list -R ANcpLua/yt-transcript --workflow=release.yml --limit 5
# Replace RUN_ID with the ID of the dispatch just started.
gh run watch RUN_ID -R ANcpLua/yt-transcript --exit-statusstores takes all, chrome, edge, or firefox. chrome takes release
(upload and submit for review) or update (upload only), for example when
the listing images change with the version: upload, swap the images in the
dashboard, then click Submit for review there. The Edge workflow step always
uses release; it has no upload-only input.
gh workflow run release.yml -R ANcpLua/yt-transcript --ref main -f stores=chrome -f chrome=updateWhat the dispatch does: lint, unit tests, listing lint, README table check,
build, zip, then for each selected store store-publish <store> release
(or chrome update). Edge ships the Chromium build with manifest.edge.json
(no tab audio, Edge has no built-in model with audio input); source is the
archive AMO requires for bundled builds.
Let the run finish: later store steps still run when an earlier store fails. Inspect each store's log and dashboard before retrying; do not rerun all stores or reuse an existing version blindly.
- If the upload was rejected before a package was accepted, resolve the cause and dispatch only that store.
- If Chrome or Edge already accepted the upload but submission failed, finish
the listing/privacy fields and submit the existing draft in its dashboard,
or run
bunx store-publish chrome publish/bunx store-publish edge publishfrom this repository with that store's credentials. These commands skip upload. - If Firefox already has the version, complete any missing source upload or review information on that version instead of creating it again.
- If a submission is in review, wait for its decision. Cancel it only when intentionally replacing that submission, not as an automatic retry.
On the owner's machine, ~/.config/vitals/edge-publish status checks the local
Edge credentials and ~/.config/vitals/edge-publish publish submits the existing
draft. Run the helper from the intended repository root; it reads that repo's
product ID from store.config.json.
Completion means recording the version and result for each store separately: upload failed, uploaded draft, submitted/in review, or live. A green workflow confirms its API operations, not eventual review approval. Expired credentials, store outages, review locks, and dashboard requirements can still interrupt a run.
Routine maintenance and the necessary follow-through are part of the user's standing authorization. Carry out the next clear step, verify its result, and report what changed. Do not stop to ask whether to check status, fix a known workflow failure, fill a field from verified repository facts, sync a renewed credential, or finish an already requested submission.
- At task start, compare recorded expiry dates with today's date and check known unresolved publishing issues. Before any store submission, refresh live status and credential checks using the runbook. Edge's credential probe does not reveal certification state; inspect its dashboard.
- Treat a key expiring within 30 days as maintenance due. Renew it through the supported flow when access allows, update the local copy and affected secrets in both extension repositories, update both configs and generated tables, and verify both repositories' authentication. Diagnose rejected credentials before rotating them; preserve the shared AMO pair as described above. Keep secret values out of logs and commits.
- Complete required dashboard fields from the current code, privacy policy, and saved listing answers. Resolve factual mismatches in the code and docs; ask only for an essential fact that those sources cannot establish.
- For a failed upload or submission, follow the partial-release rules above and continue the other stores. For an open review, retain the submission and report its current state and next check; review approval is external. Investigate a rejection and implement the justified fix before resubmitting within the authorized release. Publish completion still requires evidence for each store separately.
- Remind the user when a deadline is approaching or a blocker needs their action. Include the store, due date or review state, work already completed, and the exact remaining action. Repeat only for a meaningful change or an approaching deadline. If login, MFA, access, or an unknown required fact prevents progress, explain that specific blocker and continue independent work.
Session instructions do not run between tasks. When reminders or review checks are scheduled, use an actual automation, identify the next check, and notify on actionable changes rather than unchanged pending reviews. Keep unresolved follow-up explicit; never promise unattended monitoring unless a scheduled check exists.
Store images live in store/images and are generated:
bunx playwright test --config scripts/store-images/playwright.config.ts
renders the five 1280x800 side-panel screenshots from a local fixture, and
node scripts/store-images/promo.mjs renders the tile and marquee. Chrome and
Edge take them in their dashboards; Firefox takes them through
store-status with amo-icon-diff/amo-icon-apply for the listing icon and
amo-previews-list/amo-previews-apply for the screenshots, outside any
version review.
Dashboard slots: store icon icon-128x128.png, screenshots 1 to 5, small
promo tile tile-440x280.png, marquee marquee-1400x560.png; Edge Partner
Center additionally takes logo-300x300.png.
Store facts that shape this:
- Chrome and Edge can block new uploads during review or certification. Wait before submitting another package. Tagging is separate and only creates the GitHub release. Chrome supports intentional review cancellation, but a retry should preserve the existing submission unless replacing it.
- Chrome rejected 3.0.0 once as keyword spam ("Yellow Argon") for a run of
file-format acronyms in the description.
store-publish lintrejects such comma chains and the words bypass, unlock, circumvent. Keep any one keyword under five uses. - A trademark complaint over a platform name in the extension name was closed by renaming to Video Transcript. No platform names in user-facing text; the adapter, permissions, fixtures, and technical docs may name one when necessary.
- AMO stores the listing under the locale the add-on was created with, which
is
deeven though the text is English. The tool reads that locale from the add-on; never hardcode one, or a second translation appears and the served one stays stale. - Check the Chrome Web Store privacy form before submission. The intended
answers are kept in
store/privacy-fields.md. - Firefox reviewers rebuild from the source zip using
AMO_BUILD.md.
| What | Expires | Renew at |
|---|---|---|
| Edge Publish API key | See the generated store table above | https://partner.microsoft.com/en-us/dashboard/microsoftedge/publishapi, then update the local copy and the two Edge secrets in both repositories, update stores.edge.expires in both configs, regenerate their README tables, and run store-publish edge status |
| Chrome refresh token | none, but Google revokes it after six months without use or on a consent-screen change | https://developers.google.com/oauthplayground with the same client, then CWS_REFRESH_TOKEN in both repositories |
| AMO key pair | none | do not regenerate, see above |
Requirements: Bun 1.4 or newer as the package manager, Node 22.11 or newer as the runtime for Vite, esbuild, and the unit tests, Chrome with side-panel support. Audio transcription additionally needs Chrome 138 or newer with the on-device model available on the machine; caption discovery does not.
bun install --frozen-lockfile
bun run lint # tsc, strict
bun run test # unit tests, Vitest
bun run build # dist/ (Chrome, Edge) and packages/extension/dist-firefox/
bunx playwright test # browser suite against the unpacked build
bun run zip # build plus the four store zipsLoad dist/ as an unpacked extension for manual checks.
Toolchain decision, shared with save-media so a third extension can copy
either repository: bun installs packages (bun.lock is the only lockfile),
Vitest runs the unit tests, Vite and esbuild bundle, Playwright runs the
browser suite. Bun is not the bundler: MV3 needs Vite's and esbuild's output.
bun test is not used: it is a third runner that is neither a drop-in for
node:test nor able to run save-media's jsdom tests, and one runner across
both repositories is worth more than the dependency Vitest adds. Run tests
with bun run test, never bun test. CI (ci.yml) runs lint, unit tests,
build, listing lint, README check, and the Playwright suite on every push and
pull request.
- Zero cost: no backend, accounts, credits, paid APIs, or server.
- AI is Chrome built-in AI only. No API keys, no bundled ML runtime.
- No tracking, analytics, cookies, or telemetry. No dependency that phones home.
- React 19, Vite, Tailwind CSS 4, strict TypeScript. No
any,@ts-ignore, or double-cast escapes. chrome.storageand IndexedDB, neverlocalStorage.- Designed for a 400 px side panel.
- No
console.login shipped code. - Public text (README, listing, release notes, manifest description, UI strings): no em dashes, no emojis, no platform or browser names, claims verified first.
- Correctness and coherence over API stability. Delete superseded paths instead of adding compatibility layers; update every caller in the same change.
- Removed features stay removed: sentiment, topics, mind map, quotes, study guide, quiz, flashcards, bilingual view, speaker detection.
npm run lint,npm test,npm run build,npx playwright test.- No new dependency without a reason, no new network host, no platform name in user-facing text.
- For anything touching discovery or audio, load
dist/unpacked and walk the manual checks below.
Manual checks in real Chrome:
- Open a captioned media page, click the toolbar action. A Native pill appears without starting playback. Pause the media and repeat; the transcript still loads.
- On a page with cross-origin media sources use Inspect media sources, grant the exact origin, and confirm cues load. The fetch runs in the service worker, so page CSP cannot block it.
- On a captionless page the panel reports that no native text was found before it offers Transcribe live audio.
- Live transcription: pausing shows a resume hint, resuming delivers segments, muting suspends capture, media end finalizes, and refusal or help prose never appears as transcript text.
- Drop a speech file and check the on-device result.
- Inspect the service worker and offscreen document for errors and confirm there are no tracking or model-provider requests.
Native timed text is always attempted before audio. Audio is an explicit fallback, never the primary path.
- L0, user-granted page discovery.
chrome.action.onClickedopens the side panel for that tab and grantsactiveTab. The service worker injectssrc/content/timed-text-bridge.ts(isolated world) andsrc/content/timed-text-main.ts(main world). No manifest content scripts, no permanent host permissions. A pasted URL opens visibly in a new tab with a SCAN badge; the user clicks the action there once. A background tab never inheritsactiveTab. The side panel is enabled per tab, so it follows the tab the user invoked it on. - L1, browser runtime tracks. The bridge inspects every accessible
<video>and<audio>:TextTrackList, runtimeTextTrackobjects,VTTCuelists, child<track>elements, all roles, and playback state. Disabled subtitle tracks are switched to hidden so Chrome loads cues without rendering them. Mutations and media events rescan. - L2, timed-text resources. The main-world observer checks prior
PerformanceResourceTimingentries, futurefetchand XHR responses,<track src>,blob:anddata:tracks. Candidate detection is centralized insrc/lib/timed-text/detect.ts; the URL-marker and MIME allowlists are locked by unit tests. Text bodies are read with a 2 MiB ceiling; binary media is never decoded as text; unrelated JSON is discarded unless it holds timestamped cues. Cross-origin text is reported as metadata and fetched by the service worker after an exact-origin grant, outside page CSP. - L3, formats and manifests.
src/lib/timed-text/parse.tsnormalizes WebVTT, SRT, TTML/DFXP/IMSC text, ASS/SSA, SAMI, SBV, LRC, and common timestamped JSON.src/lib/timed-text/manifest.tsexpands HLS subtitle playlists and DASH text representations and records CEA signaling and binary text it cannot decode. Bitmap, broadcast, and MP4 timed text are detected, never claimed as decoded. - L4, optional page adapter. If generic discovery finds nothing on a
compatible watch page,
src/content/adapters/youtube.tsis injected once as a bounded fallback. It is not statically registered and never leaks platform wording into the UI. Playlist, channel, CSV, and bare-ID bulk tools ask for optional source-site access from the user's click. - L5, live audio.
src/background/transcribe/tab-capture.tsconsumes a user-grantedtabCapturestream in the offscreen document, resamples to 16 kHz mono, and prompts Chrome's on-device model in 8-second windows. The pipeline suspends on pause or mute, finalizes at media end, timestamps from media time, skips silence, discards refusal prose, and keeps tab audio audible. Availability is gated byLanguageModel.availability(...)with audio input declared; unsupported hardware fails loudly. Dropped files go throughOfflineAudioContext.decodeAudioDatainto the same window function.
Messages between the service worker, offscreen document, and side panel are
validated with zod schemas in src/lib/messages/schema.ts; the message types
are inferred from them. Content scripts only import types, so the zod runtime
never ships into page frames.
Verified on real hardware (Apple silicon, Chrome 152, 2026-09-12): the extension's own prompt transcribed a 6-second and a 20-second synthetic speech sample verbatim except for words cut by a window boundary. A warm window of 8 seconds takes about 7 seconds; the first inference after a cold start takes about 20 seconds. The model has no word timestamps, so segment timing stays window-granular.
| Capability | Status |
|---|---|
Runtime TextTrack and VTTCue |
Browser-tested while paused |
Native <track src> WebVTT |
Browser-tested while paused |
| SRT, TTML/DFXP, ASS/SSA, SAMI, SBV, LRC, timestamped JSON | Fixture-tested |
| HLS WebVTT and DASH text discovery | Fixture-tested |
| Cross-origin embedded player | Optional exact-origin permission flow |
| MP4/fMP4, CEA, bitmap, broadcast formats | Detect-only unless runtime cues exist |
| Optional authenticated page adapter | Deterministic browser fixture |
| Live and file on-device transcription | Verified on Apple silicon; Chrome only |
| Firefox | Native discovery, saving, export; no on-device audio |
Features: single-page extraction, playlist, CSV, and channel bulk extraction, history, saved transcripts with highlights, notes, and tags, summary, key points, Q&A, transcript chat, filler removal, chapter extraction, export to text and subtitle formats, cancelable AI requests, click-to-seek timestamps.
src/
background/
service-worker.ts message router, action click, install hygiene
install.ts drops permissions the manifest no longer declares
panel/ side panel (Chromium) and sidebar (Firefox)
discovery/coordinator.ts L0 to L4 orchestration and SCAN badge
transcribe/ tab capture, offscreen document, worklet
providers/ bulk and ID-only adapter
content/ injected discovery scripts, optional adapter
lib/
messages/schema.ts zod wire schemas, inferred message types
timed-text/ detect, parse, manifest
transcription/audio.ts
sidepanel/App.tsx
components/
types/ type-only views of the schemas
scripts/
build.mjs Vite + esbuild for Chrome and Firefox
package-stores.mjs the four store zips
store/
listing.md Chrome and AMO description, reviewed source
privacy-fields.md Chrome privacy form answers
store.config.json store ids, URLs, credential names
e2e/, tests/unit/ Playwright suite, unit tests
Required permissions: sidePanel, activeTab, scripting, storage,
tabCapture, offscreen. Optional HTTP(S) host access exists only so the
user can grant an embedded frame or bulk source at runtime. There is no
permanent host permission and no always-on content script. See
PRIVACY.md.
.github/workflows/credential-health.yml checks Chrome, Edge, and Firefox
credentials every six hours and through Run workflow. These are read-only
checks: nothing is uploaded or published. Vitals reads the secret metadata and
per-store authentication results without downloading secret values. A replaced
secret needs a newer successful check; results older than 24 hours are stale.
The Edge check reads a previous real publishing operation, so a 404 or server
error is never reported as successful authentication.