Skip to content

Latest commit

 

History

172 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Video Transcript

Browser extension (Chrome, Edge, Firefox) that turns the captions a media page already carries into searchable, exportable text. It inspects a page only after the user clicks the toolbar action on it. When a page has no readable timed text, Chrome's built-in on-device model can transcribe the tab's audio or a dropped file. No backend, accounts, API keys, analytics, or paid services.

This file is the operating manual. Everything an operator or a future agent needs is here or linked from here.

Stores

Store Listing Dashboard API docs Credentials (GitHub Actions secrets) Notes
Chrome Web Store ahddbfbjafmbceehebpeanpnlbaimepk dashboard docs CWS_CLIENT_ID, CWS_CLIENT_SECRET, CWS_REFRESH_TOKEN, CWS_PUBLISHER_ID Listing text, screenshots and privacy answers are edited in the dashboard; answers are kept in store/privacy-fields.md
Microsoft Edge Add-ons 069ca91d-a7cd-4bac-8224-1ee38a2d2a06 dashboard docs EDGE_API_KEY, EDGE_CLIENT_ID API key expires 2026-11-29. Chromium build with manifest.edge.json: no tabCapture or offscreen, the built-in model has no audio input there. Renew the key at https://partner.microsoft.com/en-us/dashboard/microsoftedge/publishapi and update it in both extension repos
Firefox Add-ons (AMO) video-transcript@qyl.at dashboard docs AMO_JWT_ISSUER, AMO_JWT_SECRET One API key pair per Mozilla account, shared with the other extension repo; a new key invalidates the old one everywhere. The listing text is applied from store/listing.md

The table is rendered from store.config.json by the shared store-publish tool, which save-media owns; this repository only consumes it and files issues there. Change the config, then run bunx store-publish readme --write; CI fails when the two drift. A new extension copies that file and changes the ids.

Credentials are GitHub Actions secrets in this repository. GitHub never returns their values; test those stored copies in an Actions run. The readable local copies listed below can also be used with the CLI. The same eight values are set on save-media, which shares the Chrome OAuth client, the Edge Publish API key, and the AMO key pair.

Two rules that have already cost a day each:

  • Never generate a new AMO key. Mozilla issues one key pair per account. A new pair invalidates the old one immediately and breaks publishing in both repositories. An HTTP 401 almost always means the stored value is mangled, not that the key is gone.
  • Never move a secret through the clipboard. A value read back with pbpaste is not always what was copied. Check the length before writing: the Edge API key is 40 characters, the Edge client id and product id are 36.

Where the readable copies live on the owner's machine (presence only, never print values):

Credential Location
AMO key pair macOS Keychain item AMO API (addons.mozilla.org); the issuer is the account field
Chrome OAuth client id and secret ~/.config/vitals/cws-client.json
Chrome refresh token ~/.config/vitals/cws-refresh-token.txt
Edge client id and API key ~/.config/vitals/store-secrets.env, mode 600
Register of all secret locations ~/.config/vitals/keys.json
Script that writes the values into both repositories ~/.config/vitals/set-store-secrets.sh

Before using set-store-secrets.sh, populate all eight values from their documented sources; the environment file alone does not supply the AMO pair. For an Edge-only renewal, update only the two Edge secrets in each repository.

The Chrome OAuth client is "Desktop client 2" in Google Cloud project server (uplifted-nuance-408417); its consent screen is in production.

Store status without publishing

gh workflow run store-status.yml -R ANcpLua/yt-transcript --ref main -f store=chrome
gh workflow run store-status.yml -R ANcpLua/yt-transcript --ref main -f store=firefox
gh workflow run store-status.yml -R ANcpLua/yt-transcript --ref main -f store=edge
gh workflow run store-status.yml -R ANcpLua/yt-transcript --ref main -f store=amo-listing-diff
gh run list -R ANcpLua/yt-transcript --workflow=store-status.yml --limit 1
gh run view <run id> -R ANcpLua/yt-transcript --log
  • chrome prints the published revision and, while a review is open, the submitted revision with its state.
  • firefox lists every version with its review status.
  • edge has no read endpoint. The step probes the credentials with an operation id that cannot exist (HTTP 401 means rejected, anything else means accepted), prints the lengths of the three values, and prints the days left until the API key expires. It fails once the key has expired.
  • amo-listing-diff compares the ## AMO description section of store/listing.md with the live AMO description. amo-listing-apply writes it. This is the only path that edits listing text; Chrome and Edge listing text is edited in their dashboards.

Release

Versions live in manifest.json, manifest.edge.json, manifest.firefox.json, and package.json, and must match. store-publish version refuses a mismatch and, on a tag, a tag that differs from them. One dispatch with stores=all attempts all three stores as sequential steps in one job. This is not a parallel job matrix. Store acceptance and review remain independent, so a successful submission is not proof that it is live.

# 1. bump all four version fields and add a CHANGELOG entry
bun install                       # refresh bun.lock after the version bump
bun run lint
bun run test
bun run build
bunx playwright test
bunx store-publish lint
bunx store-publish readme --check
release_version=$(bunx store-publish version)
# 2. stage the reviewed release changes, commit, push main, and wait for ci.yml
git add -p
git commit -m "Release $release_version"
git push origin main
# 3. after CI succeeds for this commit, create a new tag and push it
git tag "v$release_version"
git push origin "v$release_version"

The tag run checks and builds the four zips and creates a GitHub release. It does not submit to stores. Wait for that run to succeed. Before uploading, run the status checks above and inspect Edge's dashboard for certification state; its API probe checks credentials only. Resolve pending listing or privacy fields and wait for existing reviews that block uploads. Confirm main still contains the tested release code and version, then dispatch once:

gh workflow run release.yml -R ANcpLua/yt-transcript --ref main -f stores=all -f chrome=release
gh run list -R ANcpLua/yt-transcript --workflow=release.yml --limit 5
# Replace RUN_ID with the ID of the dispatch just started.
gh run watch RUN_ID -R ANcpLua/yt-transcript --exit-status

stores takes all, chrome, edge, or firefox. chrome takes release (upload and submit for review) or update (upload only), for example when the listing images change with the version: upload, swap the images in the dashboard, then click Submit for review there. The Edge workflow step always uses release; it has no upload-only input.

gh workflow run release.yml -R ANcpLua/yt-transcript --ref main -f stores=chrome -f chrome=update

What the dispatch does: lint, unit tests, listing lint, README table check, build, zip, then for each selected store store-publish <store> release (or chrome update). Edge ships the Chromium build with manifest.edge.json (no tab audio, Edge has no built-in model with audio input); source is the archive AMO requires for bundled builds.

Retrying a partial release

Let the run finish: later store steps still run when an earlier store fails. Inspect each store's log and dashboard before retrying; do not rerun all stores or reuse an existing version blindly.

  • If the upload was rejected before a package was accepted, resolve the cause and dispatch only that store.
  • If Chrome or Edge already accepted the upload but submission failed, finish the listing/privacy fields and submit the existing draft in its dashboard, or run bunx store-publish chrome publish / bunx store-publish edge publish from this repository with that store's credentials. These commands skip upload.
  • If Firefox already has the version, complete any missing source upload or review information on that version instead of creating it again.
  • If a submission is in review, wait for its decision. Cancel it only when intentionally replacing that submission, not as an automatic retry.

On the owner's machine, ~/.config/vitals/edge-publish status checks the local Edge credentials and ~/.config/vitals/edge-publish publish submits the existing draft. Run the helper from the intended repository root; it reads that repo's product ID from store.config.json.

Completion means recording the version and result for each store separately: upload failed, uploaded draft, submitted/in review, or live. A green workflow confirms its API operations, not eventual review approval. Expired credentials, store outages, review locks, and dashboard requirements can still interrupt a run.

Publishing maintenance

Routine maintenance and the necessary follow-through are part of the user's standing authorization. Carry out the next clear step, verify its result, and report what changed. Do not stop to ask whether to check status, fix a known workflow failure, fill a field from verified repository facts, sync a renewed credential, or finish an already requested submission.

  • At task start, compare recorded expiry dates with today's date and check known unresolved publishing issues. Before any store submission, refresh live status and credential checks using the runbook. Edge's credential probe does not reveal certification state; inspect its dashboard.
  • Treat a key expiring within 30 days as maintenance due. Renew it through the supported flow when access allows, update the local copy and affected secrets in both extension repositories, update both configs and generated tables, and verify both repositories' authentication. Diagnose rejected credentials before rotating them; preserve the shared AMO pair as described above. Keep secret values out of logs and commits.
  • Complete required dashboard fields from the current code, privacy policy, and saved listing answers. Resolve factual mismatches in the code and docs; ask only for an essential fact that those sources cannot establish.
  • For a failed upload or submission, follow the partial-release rules above and continue the other stores. For an open review, retain the submission and report its current state and next check; review approval is external. Investigate a rejection and implement the justified fix before resubmitting within the authorized release. Publish completion still requires evidence for each store separately.
  • Remind the user when a deadline is approaching or a blocker needs their action. Include the store, due date or review state, work already completed, and the exact remaining action. Repeat only for a meaningful change or an approaching deadline. If login, MFA, access, or an unknown required fact prevents progress, explain that specific blocker and continue independent work.

Session instructions do not run between tasks. When reminders or review checks are scheduled, use an actual automation, identify the next check, and notify on actionable changes rather than unchanged pending reviews. Keep unresolved follow-up explicit; never promise unattended monitoring unless a scheduled check exists.

Listing changes

Store images live in store/images and are generated: bunx playwright test --config scripts/store-images/playwright.config.ts renders the five 1280x800 side-panel screenshots from a local fixture, and node scripts/store-images/promo.mjs renders the tile and marquee. Chrome and Edge take them in their dashboards; Firefox takes them through store-status with amo-icon-diff/amo-icon-apply for the listing icon and amo-previews-list/amo-previews-apply for the screenshots, outside any version review. Dashboard slots: store icon icon-128x128.png, screenshots 1 to 5, small promo tile tile-440x280.png, marquee marquee-1400x560.png; Edge Partner Center additionally takes logo-300x300.png.

Store facts that shape this:

  • Chrome and Edge can block new uploads during review or certification. Wait before submitting another package. Tagging is separate and only creates the GitHub release. Chrome supports intentional review cancellation, but a retry should preserve the existing submission unless replacing it.
  • Chrome rejected 3.0.0 once as keyword spam ("Yellow Argon") for a run of file-format acronyms in the description. store-publish lint rejects such comma chains and the words bypass, unlock, circumvent. Keep any one keyword under five uses.
  • A trademark complaint over a platform name in the extension name was closed by renaming to Video Transcript. No platform names in user-facing text; the adapter, permissions, fixtures, and technical docs may name one when necessary.
  • AMO stores the listing under the locale the add-on was created with, which is de even though the text is English. The tool reads that locale from the add-on; never hardcode one, or a second translation appears and the served one stays stale.
  • Check the Chrome Web Store privacy form before submission. The intended answers are kept in store/privacy-fields.md.
  • Firefox reviewers rebuild from the source zip using AMO_BUILD.md.

Expiry dates

What Expires Renew at
Edge Publish API key See the generated store table above https://partner.microsoft.com/en-us/dashboard/microsoftedge/publishapi, then update the local copy and the two Edge secrets in both repositories, update stores.edge.expires in both configs, regenerate their README tables, and run store-publish edge status
Chrome refresh token none, but Google revokes it after six months without use or on a consent-screen change https://developers.google.com/oauthplayground with the same client, then CWS_REFRESH_TOKEN in both repositories
AMO key pair none do not regenerate, see above

Development

Requirements: Bun 1.4 or newer as the package manager, Node 22.11 or newer as the runtime for Vite, esbuild, and the unit tests, Chrome with side-panel support. Audio transcription additionally needs Chrome 138 or newer with the on-device model available on the machine; caption discovery does not.

bun install --frozen-lockfile
bun run lint         # tsc, strict
bun run test         # unit tests, Vitest
bun run build        # dist/ (Chrome, Edge) and packages/extension/dist-firefox/
bunx playwright test # browser suite against the unpacked build
bun run zip          # build plus the four store zips

Load dist/ as an unpacked extension for manual checks.

Toolchain decision, shared with save-media so a third extension can copy either repository: bun installs packages (bun.lock is the only lockfile), Vitest runs the unit tests, Vite and esbuild bundle, Playwright runs the browser suite. Bun is not the bundler: MV3 needs Vite's and esbuild's output. bun test is not used: it is a third runner that is neither a drop-in for node:test nor able to run save-media's jsdom tests, and one runner across both repositories is worth more than the dependency Vitest adds. Run tests with bun run test, never bun test. CI (ci.yml) runs lint, unit tests, build, listing lint, README check, and the Playwright suite on every push and pull request.

Constraints

  • Zero cost: no backend, accounts, credits, paid APIs, or server.
  • AI is Chrome built-in AI only. No API keys, no bundled ML runtime.
  • No tracking, analytics, cookies, or telemetry. No dependency that phones home.
  • React 19, Vite, Tailwind CSS 4, strict TypeScript. No any, @ts-ignore, or double-cast escapes.
  • chrome.storage and IndexedDB, never localStorage.
  • Designed for a 400 px side panel.
  • No console.log in shipped code.
  • Public text (README, listing, release notes, manifest description, UI strings): no em dashes, no emojis, no platform or browser names, claims verified first.
  • Correctness and coherence over API stability. Delete superseded paths instead of adding compatibility layers; update every caller in the same change.
  • Removed features stay removed: sentiment, topics, mind map, quotes, study guide, quiz, flashcards, bilingual view, speaker detection.

Verifying a change

  1. npm run lint, npm test, npm run build, npx playwright test.
  2. No new dependency without a reason, no new network host, no platform name in user-facing text.
  3. For anything touching discovery or audio, load dist/ unpacked and walk the manual checks below.

Manual checks in real Chrome:

  1. Open a captioned media page, click the toolbar action. A Native pill appears without starting playback. Pause the media and repeat; the transcript still loads.
  2. On a page with cross-origin media sources use Inspect media sources, grant the exact origin, and confirm cues load. The fetch runs in the service worker, so page CSP cannot block it.
  3. On a captionless page the panel reports that no native text was found before it offers Transcribe live audio.
  4. Live transcription: pausing shows a resume hint, resuming delivers segments, muting suspends capture, media end finalizes, and refusal or help prose never appears as transcript text.
  5. Drop a speech file and check the on-device result.
  6. Inspect the service worker and offscreen document for errors and confirm there are no tracking or model-provider requests.

How extraction works

Native timed text is always attempted before audio. Audio is an explicit fallback, never the primary path.

  • L0, user-granted page discovery. chrome.action.onClicked opens the side panel for that tab and grants activeTab. The service worker injects src/content/timed-text-bridge.ts (isolated world) and src/content/timed-text-main.ts (main world). No manifest content scripts, no permanent host permissions. A pasted URL opens visibly in a new tab with a SCAN badge; the user clicks the action there once. A background tab never inherits activeTab. The side panel is enabled per tab, so it follows the tab the user invoked it on.
  • L1, browser runtime tracks. The bridge inspects every accessible <video> and <audio>: TextTrackList, runtime TextTrack objects, VTTCue lists, child <track> elements, all roles, and playback state. Disabled subtitle tracks are switched to hidden so Chrome loads cues without rendering them. Mutations and media events rescan.
  • L2, timed-text resources. The main-world observer checks prior PerformanceResourceTiming entries, future fetch and XHR responses, <track src>, blob: and data: tracks. Candidate detection is centralized in src/lib/timed-text/detect.ts; the URL-marker and MIME allowlists are locked by unit tests. Text bodies are read with a 2 MiB ceiling; binary media is never decoded as text; unrelated JSON is discarded unless it holds timestamped cues. Cross-origin text is reported as metadata and fetched by the service worker after an exact-origin grant, outside page CSP.
  • L3, formats and manifests. src/lib/timed-text/parse.ts normalizes WebVTT, SRT, TTML/DFXP/IMSC text, ASS/SSA, SAMI, SBV, LRC, and common timestamped JSON. src/lib/timed-text/manifest.ts expands HLS subtitle playlists and DASH text representations and records CEA signaling and binary text it cannot decode. Bitmap, broadcast, and MP4 timed text are detected, never claimed as decoded.
  • L4, optional page adapter. If generic discovery finds nothing on a compatible watch page, src/content/adapters/youtube.ts is injected once as a bounded fallback. It is not statically registered and never leaks platform wording into the UI. Playlist, channel, CSV, and bare-ID bulk tools ask for optional source-site access from the user's click.
  • L5, live audio. src/background/transcribe/tab-capture.ts consumes a user-granted tabCapture stream in the offscreen document, resamples to 16 kHz mono, and prompts Chrome's on-device model in 8-second windows. The pipeline suspends on pause or mute, finalizes at media end, timestamps from media time, skips silence, discards refusal prose, and keeps tab audio audible. Availability is gated by LanguageModel.availability(...) with audio input declared; unsupported hardware fails loudly. Dropped files go through OfflineAudioContext.decodeAudioData into the same window function.

Messages between the service worker, offscreen document, and side panel are validated with zod schemas in src/lib/messages/schema.ts; the message types are inferred from them. Content scripts only import types, so the zod runtime never ships into page frames.

Verified on real hardware (Apple silicon, Chrome 152, 2026-09-12): the extension's own prompt transcribed a 6-second and a 20-second synthetic speech sample verbatim except for words cut by a window boundary. A warm window of 8 seconds takes about 7 seconds; the first inference after a cold start takes about 20 seconds. The model has no word timestamps, so segment timing stays window-granular.

Capabilities

Capability Status
Runtime TextTrack and VTTCue Browser-tested while paused
Native <track src> WebVTT Browser-tested while paused
SRT, TTML/DFXP, ASS/SSA, SAMI, SBV, LRC, timestamped JSON Fixture-tested
HLS WebVTT and DASH text discovery Fixture-tested
Cross-origin embedded player Optional exact-origin permission flow
MP4/fMP4, CEA, bitmap, broadcast formats Detect-only unless runtime cues exist
Optional authenticated page adapter Deterministic browser fixture
Live and file on-device transcription Verified on Apple silicon; Chrome only
Firefox Native discovery, saving, export; no on-device audio

Features: single-page extraction, playlist, CSV, and channel bulk extraction, history, saved transcripts with highlights, notes, and tags, summary, key points, Q&A, transcript chat, filler removal, chapter extraction, export to text and subtitle formats, cancelable AI requests, click-to-seek timestamps.

Layout

src/
  background/
    service-worker.ts          message router, action click, install hygiene
    install.ts                 drops permissions the manifest no longer declares
    panel/                     side panel (Chromium) and sidebar (Firefox)
    discovery/coordinator.ts   L0 to L4 orchestration and SCAN badge
    transcribe/                tab capture, offscreen document, worklet
    providers/                 bulk and ID-only adapter
  content/                     injected discovery scripts, optional adapter
  lib/
    messages/schema.ts         zod wire schemas, inferred message types
    timed-text/                detect, parse, manifest
    transcription/audio.ts
  sidepanel/App.tsx
  components/
  types/                       type-only views of the schemas
scripts/
  build.mjs                    Vite + esbuild for Chrome and Firefox
  package-stores.mjs           the four store zips
store/
  listing.md                   Chrome and AMO description, reviewed source
  privacy-fields.md            Chrome privacy form answers
store.config.json              store ids, URLs, credential names
e2e/, tests/unit/              Playwright suite, unit tests

Privacy

Required permissions: sidePanel, activeTab, scripting, storage, tabCapture, offscreen. Optional HTTP(S) host access exists only so the user can grant an embedded frame or bulk source at runtime. There is no permanent host permission and no always-on content script. See PRIVACY.md.

License

MIT

Credential health monitoring

.github/workflows/credential-health.yml checks Chrome, Edge, and Firefox credentials every six hours and through Run workflow. These are read-only checks: nothing is uploaded or published. Vitals reads the secret metadata and per-store authentication results without downloading secret values. A replaced secret needs a newer successful check; results older than 24 hours are stale. The Edge check reads a previous real publishing operation, so a 404 or server error is never reported as successful authentication.

About

MV3 side-panel extension that extracts and exports transcripts. Chrome, Edge, Firefox.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages