Overview
I’ve always loved those scenes in movies where a character uses a voice recorder to capture real-time notes, but never quite enough to buy a voice recorder. I have however, been interested in the Pebble since the initial run, and since they came back and I have adult money, I’ve acquired a Pebble Time 2.
It seems like everyone’s first project with the Pebble is to re-invent voice notes, and I’m no different.
Technical Details
Architecture
flowchart LR subgraph Watch["Pebble Watch"] A[Voice Note<br/>Recording] H[Sync Action] end subgraph Phone["iPhone"] B["Dictation API<br/>(Local Transcription)"] C["Review Screen<br/>(optional)"] F{Server<br/>Available?} G[(Pending Queue<br/>localStorage)] end subgraph Mac["MacBook Server"] D[Webhook<br/>Endpoint] E[(Obsidian Vault<br/>Inbox)] end A -->|"audio"| B B -->|"transcript"| C B -.->|"skip review"| F C -->|"confirmed text"| F F -- yes --> D F -- no --> G H -.->|"trigger flush"| G G -->|"flush pending"| D D -->|"writes note"| E
Use
The first priority here is easy voice recording. To that end, the app is built with a pair of interaction modes:
Quick Launch
When started with the APP_LAUNCH_QUICK_LAUNCH launch reason, Captain's Log immediately begins a dictation session. When the dictation session is ended (either by button press or timeout1), the transcript is optionally presented to the user for review. This review behavior is configurable, but transcription is off often enough that I’ve left it on.
When the transcript is approved (or automatically, if review is disabled), the text is sent to a configured webhook URL by PebbleKit JS running phone-side. If that request fails for any reason, the transcript is stored in localStorage, and a count of pending transcripts is sent back to the watch to display on the App Glance.
Log Management
When started by any other launch reason, the app opens into an action menu, allowing the user to:
- trigger bulk sync of pending transcripts to the webhook
- review individual transcripts, allowing for sync or deletion
- add a new entry
Server-Side
The webhook I have configured points to a simple Python webserver running on my home machine, and only accessible on my local network2.
All that webserver does is accept transcript text and send it to an inbox/ folder in my local Obsidian vault for me to clean up and integrate into the vault.
Issues and Next Steps
Local-first Recording Transcription
The Pebble only exposes a dictation API, rather than any kind of raw audio for transport. The dictation API farms out recordings to the phone app for processing either using a local model (on the phone) or a cloud provider.
Frankly, local transcription on the iPhone isn’t great, and I don’t want to use the WisprFlow cloud provider. The Pebble app uses parakeet-tdt-0.6b-v3 on my iPhone, and it is both slow and inaccurate. I’d love for raw recordings to be made available such that I could pass a recording to the server on my MacBook to transcribe using a beefier local model, but the Pebble doesn’t expose that functionality yet.
This PR looks like it’ll add mic API support to the OS, enabling raw recordings to be shipped to the server for transcription; I’ll be watching that before making further updates.
More Actions
Once I can more accurately process voice recordings, there’s an idea brewing for a categorization scheme that routes transcripts to different locations and turns the log into more of a voice assistant. This categorization scheme would be based on the first word or phrase of the transcript, and might look like:
Note: routed to ObsidianCommand, <project>: a prompt routed to a coding agent running in a sandboxTask, <project>: a to-do added to a running project file, also in Obsidian
This whole scheme, however, is going to be blocked behind better dictation, as I’m not loving re-recording notes for failed transcription on that category.
Open Sourcing
I haven’t opened the code here, for the primary reason that it’s heavily vibe-coded and I want to take a cleanup pass before doing so. Once that’s done, I’ll open the repo and link here.