Merge main to validate received-change readiness with Commonlib 0.1.34

This commit is contained in:
vorotamoroz
2026-09-30 16:17:19 +00:00
13 changed files with 546 additions and 27 deletions
+2 -1
View File
@@ -98,7 +98,8 @@ Synchronisation status is shown in the status bar with the following icons.
- 🛫 Pending read storage processes
- 📬 Batched read storage processes
- ⚙️ Working or pending storage processes for hidden files
- 🧩 Waiting chunks
- 🛄 Pending initial on-demand chunk requests
- 🔁 Pending chunk retries, including retry delays
- 🔌 Working customisation items (configuration, snippets, and plug-ins)
To prevent file and database corruption, please avoid closing Obsidian until all progress indicators have disappeared as much as possible (although the plug-in will attempt to resume if interrupted). This is especially important if you have deleted or renamed files.
+2 -2
View File
@@ -43,7 +43,7 @@ Applying downloaded document changes enters one local-application boundary from
Manual P2P commands which bypass `ReplicationService` enter the broad boundary. Direct P2P pull and push entry points are therefore both protected as finite remote work, covering the Obsidian panes, CLI, and Webapp. A pull or bidirectional synchronisation also enters the narrower finite-replication boundary because it can place documents in the local database. A push-only request remains broad-only: it cannot satisfy a local missing-chunk read and must not present itself as a delivery source. Automatic synchronisation on peer discovery, a pull requested by a remote peer, and a watched pull following a peer progress notification enter both boundaries because each can deliver local documents. A normal P2P peer-selection dialogue represents one broad finite session: it remains inside the boundary while waiting for a peer and while the person may perform repeated synchronisations, then settles when the dialogue closes and any in-flight synchronisation has finished. Closing without synchronising returns a failed result and releases the boundary. The 'Start Sync & Close' action completes its synchronisation before closing. This deliberately protects peer discovery and selection, because display sleep can interrupt discovery or connection establishment and require the person to start detection again. It may therefore retain a Wake Lock longer than the network transfer alone. A transfer performed inside that session temporarily adds a nested activity; the count remains a logical-operation count rather than a connection total.
`ChunkFetcher` enters the broad boundary synchronously when it accepts newly missing chunk identifiers, but it does not increment the finite-replication count. A typed per-identifier claim keeps the broad boundary active through queue waiting, interval throttling, `fetchRemoteChunks`, validation, local persistence, and terminal event delivery. Duplicate requests share the existing claim. Explicit absence, failure, cancellation, or a conservative five-minute period without fetcher-observable progress settles the affected claim. The five-minute value is only a last-resort leak fuse: it prevents a never-settling integration from retaining the per-identifier claim and waiter indefinitely and, once the activity runner has entered the claim task, lets the associated Wake Lock, lifecycle deferral, and indicator finish. It is not an arrival estimate, proof of remote absence, or a transport deadline. This scope is defined in detail by the chunk-arrival-quiescence ADR.
`ChunkFetcher` enters the broad boundary synchronously when it accepts newly missing chunk identifiers, but it does not increment the finite-replication count. A typed per-identifier claim keeps the broad boundary active through queue waiting, interval throttling, `fetchRemoteChunks`, validation, local persistence, and terminal event delivery. Duplicate requests share the existing claim. A first successful omission schedules a two-second retry; further omissions increase the per-identifier delay up to ten seconds while finite replication remains active. Backoff releases physical concurrency without releasing the bounded task. Observed finite completion interrupts backoff for a local recheck and, if needed, a final remote probe. A current final successful omission settles the claim; failure, cancellation, or a conservative five-minute period without fetcher-observable progress also settles it. The five-minute value is only an inactivity leak fuse, not a total retry budget: it prevents a never-settling integration from retaining the per-identifier claim and waiter indefinitely and, once the activity runner has entered the claim task, lets the associated Wake Lock, lifecycle deferral, and indicator finish. It is not an arrival estimate, proof of remote absence, or a transport deadline. This scope is defined in detail by the chunk-arrival-quiescence ADR.
Rebuild operations use the same boundary at their destructive or remote phase:
@@ -105,7 +105,7 @@ Unit tests cover:
- direct P2P pull and push entry points entering the broad boundary, while only pull and bidirectional operations enter the finite-delivery boundary;
- automatic synchronisation on peer discovery, remote pull requests, and watched peer progress entering the boundary;
- P2P peer-selection sessions settling on close, including cancellation, repeated synchronisation, and a close during in-flight work;
- remote chunk fetching remaining inside the shared boundary from synchronous queue acceptance through local persistence and terminal notification;
- remote chunk fetching remaining inside the shared boundary from synchronous queue acceptance through the delayed retry of omitted identifiers, local persistence, and terminal notification;
- missing-chunk waiters rechecking local storage when observed per-identifier claims and finite replication have settled;
- the finite-replication count excluding other bounded work;
- replicated document application sharing one local boundary, settling after the final recovery snapshot, and releasing around processing suspension;
+31 -7
View File
@@ -12,6 +12,8 @@ Before this decision, a missing-chunk waiter used a fixed 5-second or 30-second
Finite replication provides a stronger boundary than elapsed wall-clock time. In the absence of an error, a finite replication does not complete until it has reached the latest sequence in its scope. Once it completes, that replication cannot deliver another chunk. An on-demand fetch is a separate finite delivery path and needs its own per-identifier boundary.
A successful CouchDB lookup which omits a requested chunk is a point-in-time observation. Metadata and Chunk documents can become visible separately, so the first such result immediately after Metadata arrives does not prove that a second lookup will return the same result. The on-demand fetch lifecycle therefore needs to distinguish an initial omission from its terminal result.
## Decision
Wait only for a delivery lifecycle which is observable when the local miss is handled. Do not guess how long an unobserved producer might take.
@@ -29,7 +31,7 @@ The waiting layer follows these rules:
2. Dispatch `missingChunks` synchronously when direct fetch is permitted. `ChunkFetcher` must claim accepted identifiers before dispatch returns, closing the scheduling gap without a timer.
3. If a matching claim or finite replication is active, wait for that observable producer.
4. A valid chunk arrival resolves the waiter immediately.
5. An explicit remote-missing result resolves it as missing immediately.
5. A terminal explicit remote-missing result resolves it as missing immediately.
6. Once every observed producer completes, read the requested identifier from the local database once more, bypassing the cache.
7. Return the rechecked chunk, or return unavailable. Do not add a fixed grace period after the producer has stopped.
8. If no producer is observable after synchronous dispatch, return unavailable immediately. There is no operation for a duration to represent.
@@ -46,17 +48,24 @@ A successful finite replication completion is the authoritative ‘latest’ bou
- configured interval throttling;
- entry into the injected bounded-activity runner;
- `fetchRemoteChunks`;
- scheduled retries for identifiers omitted from successful responses;
- response validation;
- local database persistence; and
- fetched or missing event delivery.
The claim settles on every terminal path, including explicit absence, no active replicator, rejection, invalid results, destruction, and cancellation. Its completion Promise is the task passed to the bounded `chunk-fetch` activity, keeping Wake Lock, application lifecycle deferral, the remote-work indicator, and missing-chunk delivery aligned.
When the first successful response omits an identifier, `ChunkFetcher` keeps that identifier claimed and schedules a retry after two seconds, even if no finite replication is active. Identifiers present in a partial response are persisted and settled immediately. While finite replication remains active, subsequent omissions schedule per-identifier delays of four, six, eight, and then ten seconds, capped at ten seconds. Backoff returns the physical request slot to the queue without releasing the logical claim. Eligible retries can share a batch with new identifiers, and the queue wakes without requiring a new missing-chunk event.
When the observed finite count falls to zero, the fetcher interrupts any remaining backoff. It rechecks local persistence and, if the identifier is still absent, makes one final remote probe, subject to the configured request interval and concurrency. An already-running lookup cannot overlap another lookup for that identifier. If it began before the latest finite completion, an omitted identifier still needs a post-completion probe. A new finite operation which ends during that probe similarly makes its negative result stale. A successful omission from a current final probe emits `missingChunkRemote` and settles the current read; it does not leave a retry pending for a future synchronisation.
The retry policy observes only finite replication, not the claim itself or the broader bounded-activity count. It does not treat the continuous live channel as a finite producer. There is no fixed total duration or attempt limit while finite replication continues; the delay cap controls request frequency, not overall lifetime. No retry state is persisted, and transport errors do not enter this missing-result retry policy.
The claim settles on every terminal path, including explicit absence after the retry, no active Replicator, rejection, invalid results, destruction, and cancellation. Its completion Promise is the task passed to the bounded `chunk-fetch` activity, keeping Wake Lock, application lifecycle deferral, the remote-work indicator, and missing-chunk delivery aligned.
### Five-minute leak fuse
An accepted on-demand claim has a separate five-minute inactivity fuse. This is a last-resort leak safety valve, not a chunk-arrival budget or a remote-request timeout.
Its purpose is to prevent a faulty integration, a never-settling Promise, or a stalled transport from retaining logical ownership indefinitely. When it fires, the coordinator releases the per-identifier claim and its waiter. If the bounded activity callback has been entered, resolving the claim also allows the associated Wake Lock, application-lifecycle deferral, and remote-work indicator to be released. `ChunkFetcher` refreshes the fuse only at observable progress points, such as entering the activity boundary, beginning and completing throttling or transfer, and completing persistence.
Its purpose is to prevent a faulty integration, a never-settling Promise, or a stalled transport from retaining logical ownership indefinitely. When it fires, the coordinator releases the per-identifier claim and its waiter. If the bounded activity callback has been entered, resolving the claim also allows the associated Wake Lock, application-lifecycle deferral, and remote-work indicator to be released. `ChunkFetcher` refreshes the fuse at observable progress points, including entry into the activity boundary, request scheduling, transfer, and persistence. Repeated successful missing responses are progress for this fuse, so it is not a five-minute total limit on retries during active finite replication.
Five minutes is deliberately a conservative operational limit, not a value derived from a network protocol, a benchmark, or evidence that a missing chunk will arrive within that period. Firing the fuse neither proves remote absence nor makes the underlying request safe to abort. The current `fetchRemoteChunks` contract has no `AbortSignal`, so a physical request may still complete after its logical claim has been released. A future cancellable transport contract should add transport-specific deadlines and explicit cancellation without changing the lifecycle-based wait rule.
@@ -64,7 +73,7 @@ Five minutes is deliberately a conservative operational limit, not a value deriv
The unbounded live channel is not a quiescence gate because it has no natural end. Its initial pull-only catch-up is finite, however, and must enter `runFiniteReplicationActivity`. This includes the one-shot parameter fallback chain: every retry remains within the catch-up boundary until it succeeds or stops. If continuous replication later restarts with adjusted parameters, the new initial catch-up enters a new finite boundary.
Once the live channel has begun, a chunk delivered through it still resolves an existing waiter immediately, but the channel itself does not keep a new waiter open. CouchDB on-demand fetching supplies its own per-identifier claim. If a future defect demonstrates a delivery race inside a live batch, that batch lifecycle should be exposed explicitly rather than approximated with another elapsed delay.
Once the live channel has begun, a chunk delivered through it still resolves an existing waiter immediately, but the channel itself does not keep a new waiter open. CouchDB on-demand fetching supplies its own per-identifier claim and at least one delayed follow-up lookup. Further retries require active finite replication; the live channel does not extend them. If another future defect requires waiting for a live batch itself, that batch lifecycle should be exposed explicitly rather than approximated with another elapsed delay.
## Ownership
@@ -80,7 +89,7 @@ Once the live channel has begun, a chunk delivered through it still resolves an
- Treat a positive deprecated `timeout` only as source-compatible opt-in to lifecycle waiting. Its numeric value no longer represents an arrival duration.
- Preserve `preventRemoteRequest`: no on-demand request is dispatched, although an already-active finite replication may satisfy the waiter.
- Preserve Promise sharing for concurrent reads of the same chunk identifier.
- Preserve immediate explicit remote-missing results.
- Preserve immediate waiter resolution when `ChunkFetcher` emits a terminal explicit remote-missing result.
- Do not change which remote types support direct on-demand fetching.
## Historical Evidence and Scope
@@ -90,6 +99,7 @@ This decision addresses the lifecycle-race class rather than treating every ‘L
- [Issue #166](https://github.com/vrtmrz/obsidian-livesync/issues/166) contained logs where chunk collection failed shortly before related chunk writes appeared. It is evidence for the timing class, although that issue's hidden-file start-up path was repaired separately and is not claimed as a direct regression test here.
- The 2021 timing fixes in [commit `39e2eab0`](https://github.com/vrtmrz/obsidian-livesync/commit/39e2eab0238d9c37e3653cdec884cbeed543fc23) and the extended leaf timeout in [commit `9facb577`](https://github.com/vrtmrz/obsidian-livesync/commit/9facb577601d8aceff7df547cd2a6f9357fdaa29) show that elapsed timeout values have historically been used to absorb the same ordering uncertainty. They do not provide a protocol basis for retaining 5-second or 30-second delays.
- Replication pacing introduced by [commit `8d66c372`](https://github.com/vrtmrz/obsidian-livesync/commit/8d66c372e15c43a2de84a223c6385077b7724eec) and commonlib [commit `051b50c`](https://github.com/vrtmrz/livesync-commonlib/commit/051b50ca38ec4c05a11e8216ac259b4488b825f0) is a direct precedent for preventing replication progress from outrunning chunk collection. The present design expresses that dependency as an explicit lifecycle and completion recheck.
- [Issue #1224](https://github.com/vrtmrz/obsidian-livesync/issues/1224) reports repeatable burst edits where the receiving device observes Metadata, finds its Chunk absent in the first direct lookup, and then receives the Chunk after the file read has already failed. Retrying addresses that transient ordering case. It does not claim to reconstruct a Chunk which was never uploaded.
- [Issue #505](https://github.com/vrtmrz/obsidian-livesync/issues/505) was traced to chunks which were genuinely absent after the former bulk-send option broke the chunks-before-metadata guarantee. Waiting cannot recreate missing data, so this decision does not claim to fix it.
- [Issue #771](https://github.com/vrtmrz/obsidian-livesync/issues/771) and [Issue #986](https://github.com/vrtmrz/obsidian-livesync/issues/986) contain ambiguous or version-dependent `Load failed` reports. They remain unclaimed until the original writer and database state can be reproduced.
@@ -113,6 +123,8 @@ The broad count includes operations which cannot provide the requested chunk. Us
An unobserved producer has no defined start, progress, or completion semantics. A timer would therefore be a guess rather than a safety property. Relevant delivery paths must claim their work synchronously or expose a finite replication boundary; otherwise the read returns unavailable.
The on-demand backoff is not such a fallback timer. `ChunkFetcher` retains the identifier claim throughout each delay and owns the lookup it schedules. Once finite activity ends, the fetcher cuts the delay short and uses a current final probe rather than waiting for an unobserved future producer.
### Remove every timer
The arrival wait has no elapsed timer, but an implementation fault can leave a delivery claim unresolved forever. The five-minute inactivity fuse bounds that leaked logical state without being used as a successful delivery condition.
@@ -125,6 +137,15 @@ Unit tests use deterministic clocks and deferred Promises to cover:
- successful finite completion causing a cache-bypassing local database recheck;
- immediate unavailability when no producer is observable;
- per-identifier claims covering queueing, throttling, remote fetch, validation, persistence, and event delivery;
- an autonomous two-second retry after a first successful omission, retaining the claim and bounded remote activity;
- per-identifier backoff capped at ten seconds, with physical concurrency released during each delay;
- mixed retry stages sharing batches without resetting their individual delays or starving eligible retries;
- finite completion interrupting backoff, local persistence avoiding the final request, and pre-completion lookups requiring a current final probe;
- partial results settling available identifiers and retrying only absent identifiers;
- partial-batch completion and expired-request responses preserving replacement claims;
- concurrent request starts respecting the configured interval after retry and throttle waits;
- a final successful omission producing the terminal remote-missing result;
- disjoint initial and retry counts, aggregate ownership across fetchers, and timer cleanup;
- explicit missing, no-replicator, rejection, invalid response, cancellation, runner rejection, and teardown paths;
- overlapping claims and finite replications;
- a runner which never enters the task and a request which never settles;
@@ -132,13 +153,16 @@ Unit tests use deterministic clocks and deferred Promises to cover:
- continuous replication's finite initial catch-up, including its parameter fallback path; and
- the setting and replicator decision matrix.
Integration-style unit tests exercise `LayeredChunkManager`, `ChunkFetcher`, a memory-backed PouchDB database, and a deferred fake replicator together. A real Obsidian test is not required because the change remains behind the existing database, service, and event boundaries and does not alter platform UI or an adapter contract.
Integration-style unit tests exercise `LayeredChunkManager`, `ChunkFetcher`, a memory-backed PouchDB database, and a deferred fake replicator together. The downstream `chunk-fetch-retry` scenario exercises real CouchDB lookup results, local persistence, Vault reflection, and the actual status bar in Obsidian. Deterministic clock tests remain responsible for the exact backoff schedule and overlapping completion races; the real-runtime scenario does not substitute HTTP responses or synthesise finite-activity counts.
## Consequences
- A healthy finite replication or on-demand request no longer loses a race against an unrelated wall-clock estimate.
- A Chunk omitted from the first successful CouchDB lookup receives a follow-up lookup while the same delivery claim remains active, and further retries while finite replication continues.
- Successful finite replication completion provides a precise latest boundary for missing-chunk reads.
- The local recheck closes event-delivery and cache timing gaps without extending the wait after completion.
- Reads no longer pause for 5 or 30 seconds when no observable operation can deliver the chunk.
- Reads no longer pause for the former 5-second or 30-second arrival budgets. Backoff schedules observable requests rather than setting an elapsed arrival deadline.
- With no finite replication active, a genuinely absent remote Chunk takes one additional lookup after two seconds. Active finite replication can extend that lifetime; its completion expedites the final probe.
- The status bar separates initial pending identifiers (`🛄`) from retrying identifiers (`🔁`). Their sum remains the internal pending count used for restart deferral.
- Relevant producers must expose a lifecycle and must continue to prove cleanup on every exceptional path.
- The five-minute fuse bounds leaked logical activity, but it neither establishes remote absence nor cancels a physical request.
@@ -53,7 +53,9 @@ For CouchDB, `readChunksOnline` changes what normal replication includes, not th
| `true` | No | Primary chunk delivery after metadata arrives. |
| `false` | Yes | Recovery fallback for a chunk which is unexpectedly unavailable locally. |
`concurrencyOfReadChunksOnline` and `minimumIntervalOfReadChunksOnline` affect only the scheduling of CouchDB on-demand requests. They do not change whether a request may be dispatched or which lifecycle a reader observes. Accepted identifiers remain claimed while they wait for a concurrency slot and while the configured interval is applied. A minimum interval of five minutes or more is an exceptional value: the inactivity fuse may release the logical claim before that deliberate pause completes. This safety precedence does not abort the delayed physical request.
`concurrencyOfReadChunksOnline` and `minimumIntervalOfReadChunksOnline` affect only the scheduling of CouchDB on-demand requests. They do not change whether a request may be dispatched or which lifecycle a reader observes. Accepted identifiers remain claimed while they wait for a concurrency slot and while the configured interval is applied. Backoff releases the physical concurrency slot but retains logical ownership. Eligible retries and new identifiers can share a batch of up to 100 identifiers. An owned timer wakes the queue at the earliest eligible retry without requiring a new event. New arrivals do not reset existing retry times, and eligible older work precedes newer queued work.
After every wait or asynchronous local recheck, the fetcher checks the configured interval against the latest shared request time and reserves its start synchronously. A minimum interval of five minutes or more is an exceptional value: the inactivity fuse may release the logical claim before that deliberate pause completes. An expired claim is not dispatched after the pause; an already-running physical request is not aborted by the fuse.
## Wait State Machine
@@ -62,11 +64,12 @@ For CouchDB, `readChunksOnline` changes what normal replication includes, not th
3. Register one shared waiter per missing identifier.
4. If policy permits direct fetch, emit `missingChunks`. `ChunkFetcher` synchronously creates the per-identifier claim before the event dispatch returns.
5. Observe both the matching claim and `finiteReplicationActivityCount`.
6. Resolve immediately if a valid chunk or explicit remote-missing event arrives.
7. If an observed producer remains active, do not charge elapsed time against an arrival budget.
8. When all observed producers end, bypass the cache and read the identifiers from the local database once.
9. Return the rechecked chunk, or return unavailable. Do not add another fixed grace after the authoritative boundary.
10. If no producer is observable after synchronous dispatch, return unavailable immediately.
6. Resolve immediately if a valid chunk or terminal explicit remote-missing event arrives.
7. If a successful direct-fetch response omits an identifier, apply the per-identifier retry policy below. Settle identifiers present in a partial result and retry only those which remain absent.
8. If an observed producer remains active, do not charge elapsed time against a general arrival budget.
9. When all observed producers end, bypass the cache and read the identifiers from the local database once.
10. Return the rechecked chunk, or return unavailable. Do not add another fixed grace after the authoritative boundary.
11. If no producer is observable after synchronous dispatch, return unavailable immediately.
If new relevant activity starts while the final database recheck is pending, that result becomes stale. The waiter remains active until the newer producer completes and a current recheck finishes.
@@ -84,18 +87,30 @@ The continuous live channel is intentionally excluded because it has no completi
## Meaning of an On-demand Claim
An accepted identifier remains claimed from synchronous queue acceptance through throttling, physical fetch, validation, local persistence, and terminal event delivery. The claim is identifier-scoped because a global remote-work count cannot say whether unrelated work can provide this chunk.
An accepted identifier remains claimed from synchronous queue acceptance through throttling, physical fetch, validation, local persistence, retry delays, and terminal event delivery. The claim is identifier-scoped because a global remote-work count cannot say whether unrelated work can provide this chunk.
The claim finishes when the fetcher has recorded an outcome for the identifier. A transport error, missing active replicator, or invalid result releases the claim without emitting an explicit remote-missing result unless the remote actually supplied that information.
The claim finishes when the fetcher has recorded an outcome for the identifier. A successful first omission schedules a two-second retry even if finite replication is already inactive. While finite replication remains active, each subsequent omission increases that identifier's delay by two seconds, up to ten seconds: `2, 4, 6, 8, 10, 10, ...`. Joining a different batch does not reset its retry stage. Each retry first bypasses the local cache and disables further remote dispatch and delivery waiting for its local recheck.
When the observed finite count falls to zero, remaining backoff is interrupted. The fetcher rechecks local persistence and makes a final remote probe only for still-absent identifiers, respecting concurrency and minimum request spacing. A pre-completion in-flight lookup cannot count as that final probe: if it omits the identifier, a subsequent post-completion lookup is needed, without overlapping requests for the same identifier. If another finite operation ends during the final probe, its negative result is stale too. The current final successful omission emits the terminal explicit remote-missing result and settles the current read. No retry remains for a future synchronisation.
A transport error, missing active Replicator, or invalid result releases the claim according to its existing terminal path rather than entering this missing-result retry. The retry gate is finite replication alone: neither the fetcher's own claim nor broader bounded remote work keeps it alive. There is no attempt limit or absolute elapsed limit while finite replication continues. All retry state is in memory; destruction and inactivity expiry remove queued work and its timer.
Each request retains the identity of the claims it accepted. If another read claims the same identifier after an earlier claim settles or expires, the earlier request cannot release the replacement claim, refresh its fuse, or report unavailability for it. This also applies when a partial batch has already settled one identifier but is still retrying another.
The status indicators count two disjoint sets of pending on-demand Chunk identifiers. `🛄` shows identifiers with no successful missing response yet; `🔁` shows identifiers omitted at least once, including backoff, retry requests, and the final probe. Thus `🛄3 🔁2` means five pending identifiers. The atomic `chunkFetchCounts` snapshot provides this classification, while `collectingChunks` retains their total for the existing restart-deferral check. Moving between categories does not change that total.
Each fetcher contributes its unique accepted identifiers from queueing through terminal delivery. Repeated requests and retry attempts do not increase the count. Settled, expired, and destroyed claims leave the count; one fetcher's teardown preserves another fetcher's contribution. These are not counts of replication connections or all missing Chunks. Zero means that no on-demand claims remain, not that every Chunk was retrieved successfully.
## Meaning of the Five-minute Value
The five-minute value is an inactivity leak fuse for an accepted on-demand claim. It is the only elapsed duration in this state machine, and it is not a normal terminal condition.
The five-minute value is an inactivity leak fuse for an accepted on-demand claim. It is not a normal terminal condition and is distinct from the backoff which schedules identified follow-up lookups.
The fuse bounds retention if a faulty activity runner never enters its task, a Promise never settles, or a transport stops making observable progress. It prevents the per-identifier claim and waiter from remaining live forever. Once the bounded activity callback has entered, releasing the claim also allows Wake Lock, application-lifecycle deferral, and the remote-work indicator associated with that callback to finish. Observable progress rearms the fuse.
Five minutes is a conservative operational ceiling rather than a measured chunk-arrival expectation. It must not be used to infer that the remote lacks a chunk, and it does not abort the physical request. `fetchRemoteChunks` does not yet accept an `AbortSignal`, so the request may complete after the logical state has been released. Transport cancellation and transport-specific deadlines are separate future work.
Backoff does not resolve the waiter by elapsed time or prove that another producer will deliver the Chunk. Successful missing responses refresh the inactivity fuse, so retries can continue beyond five minutes while finite replication remains active. This is an intentional distinction between inactivity protection and a total lifetime limit.
The old 5-second and 30-second constants remain exported for source compatibility only. A positive deprecated `ChunkReadOptions.timeout` opts into lifecycle waiting, but its numeric value is ignored. Zero or a negative value still requests an immediate result. New code uses `waitForDelivery` explicitly.
## Test Obligations
@@ -110,6 +125,17 @@ Changes to this behaviour must keep automated coverage for:
- overlapping finite operations and overlapping per-identifier claims;
- activity restarting while a local recheck is pending;
- direct fetch queueing, throttling, persistence, and terminal notification;
- an autonomous two-second retry after a first successful omission, with bounded remote activity retained throughout;
- per-identifier `2, 4, 6, 8, 10, 10, ...` backoff while finite replication is active;
- backoff releasing physical concurrency, mixed-stage batching, and eligible retries not being starved by new work;
- expiry of the earliest queued claim preserving the scheduled retry for later identifiers;
- finite completion interrupting backoff and pre-completion in-flight requests requiring a current final probe;
- local persistence avoiding an unnecessary retry, including completion during request-interval throttling;
- partial fetch results settling available identifiers and retrying only absent identifiers;
- partial-batch completion and expired-request responses preserving replacement claims;
- disjoint initial and retry counts retaining queued identifiers without duplicates and releasing only settled ownership;
- concurrent request starts respecting the configured interval after retry and throttle waits;
- terminal unavailability after a current final successful omission;
- explicit remote absence versus transport or replicator failure;
- runner rejection, cancellation, teardown, and an operation which never enters its task;
- leak-fuse refresh at observable progress points; and
+4 -4
View File
@@ -23,7 +23,7 @@
"@smithy/types": "^4.14.3",
"@smithy/util-retry": "^4.4.5",
"@vrtmrz/browser-ui-kit": "0.1.0",
"@vrtmrz/livesync-commonlib": "0.1.33",
"@vrtmrz/livesync-commonlib": "0.1.34",
"@vrtmrz/obsidian-plugin-kit": "0.1.4",
"@vrtmrz/ui-interactions": "0.1.2",
"diff-match-patch": "^1.0.5",
@@ -4567,9 +4567,9 @@
}
},
"node_modules/@vrtmrz/livesync-commonlib": {
"version": "0.1.33",
"resolved": "https://registry.npmjs.org/@vrtmrz/livesync-commonlib/-/livesync-commonlib-0.1.33.tgz",
"integrity": "sha512-jziEUYCrGty3btR8pj/nkQrfLIGsQBbMug9opN8vCwAZmAznS0Zuh/l1TJpoKl8nqijZT1AJ4w5jkgU9pEyz4Q==",
"version": "0.1.34",
"resolved": "https://registry.npmjs.org/@vrtmrz/livesync-commonlib/-/livesync-commonlib-0.1.34.tgz",
"integrity": "sha512-EeVpeFg3W43cpN+x0BZvFj9fN+eQFil9vohvmLLV1z/8bCTT0Hsq/bfRV59oL0qFNggPetkwyCZUusdfjajNOw==",
"license": "MIT",
"dependencies": {
"@aws-sdk/client-s3": "^3.808.0",
+2 -1
View File
@@ -66,6 +66,7 @@
"test:e2e:obsidian:p2p-pane": "tsx test/e2e-obsidian/scripts/p2p-pane.ts",
"test:e2e:obsidian:vault-reflection": "tsx test/e2e-obsidian/scripts/vault-reflection.ts",
"test:e2e:obsidian:couchdb-upload": "tsx test/e2e-obsidian/scripts/couchdb-upload.ts",
"test:e2e:obsidian:chunk-fetch-retry": "tsx test/e2e-obsidian/scripts/chunk-fetch-retry.ts",
"test:e2e:obsidian:tweak-compatibility": "tsx test/e2e-obsidian/scripts/tweak-compatibility.ts",
"test:e2e:obsidian:couchdb-manual-setup-workflow": "tsx test/e2e-obsidian/scripts/couchdb-manual-setup-workflow.ts",
"test:e2e:obsidian:cli-to-obsidian-sync": "tsx test/e2e-obsidian/scripts/cli-to-obsidian-sync.ts",
@@ -186,7 +187,7 @@
"@smithy/types": "^4.14.3",
"@smithy/util-retry": "^4.4.5",
"@vrtmrz/browser-ui-kit": "0.1.0",
"@vrtmrz/livesync-commonlib": "0.1.33",
"@vrtmrz/livesync-commonlib": "0.1.34",
"@vrtmrz/obsidian-plugin-kit": "0.1.4",
"@vrtmrz/ui-interactions": "0.1.2",
"diff-match-patch": "^1.0.5",
+7 -3
View File
@@ -10,7 +10,7 @@ import {
import { scheduleTask } from "octagonal-wheels/concurrency/task";
import { fireAndForget, isDirty, throttle } from "@vrtmrz/livesync-commonlib/compat/common/utils";
import {
collectingChunks,
chunkFetchCounts,
pluginScanningCount,
hiddenFilesEventCount,
hiddenFilesProcessingCount,
@@ -36,7 +36,11 @@ import {
formatRemoteActivityStatusLabel,
getTrackedRequestCount,
} from "./RemoteActivityStatus.ts";
import { createMinimumVisibleActivityCount, createPaddedCounterLabel } from "./StatusBarDisplay.ts";
import {
createChunkFetchCounterLabel,
createMinimumVisibleActivityCount,
createPaddedCounterLabel,
} from "./StatusBarDisplay.ts";
import type { LiveSyncCore } from "@/main.ts";
import { LiveSyncError } from "@vrtmrz/livesync-commonlib/compat/common/LSError";
import { isValidPath } from "@/common/utils.ts";
@@ -140,7 +144,7 @@ export class ModuleLog extends AbstractObsidianModule {
const labelStorageCount = registerDisplay(
createPaddedCounterLabel(this.services.replication.storageApplyingCount, `💾`)
);
const labelChunkCount = registerDisplay(createPaddedCounterLabel(collectingChunks, `🧩`));
const labelChunkCount = registerDisplay(createChunkFetchCounterLabel(chunkFetchCounts));
const labelPluginScanCount = registerDisplay(createPaddedCounterLabel(pluginScanningCount, `🔌`));
const labelConflictProcessCount = registerDisplay(
createPaddedCounterLabel(this.services.conflict.conflictProcessQueueCount, `🔩`)
+47
View File
@@ -131,3 +131,50 @@ export function createPaddedCounterLabel(
source.offChanged(update);
});
}
/**
* Displays the disjoint initial and retry chunk-fetch counts with the same
* padding and inactive linger behaviour as the other status counters.
*/
export function createChunkFetchCounterLabel(
source: ReactiveValue<{ initial: number; retrying: number }>
): DisposableReactiveValue<string> {
const initialCount = reactiveSource(0);
const retryingCount = reactiveSource(0);
const initialLabel = createPaddedCounterLabel(initialCount, "🛄");
const retryingLabel = createPaddedCounterLabel(retryingCount, "🔁");
const formatted = reactiveSource(`${initialLabel.value}${retryingLabel.value}`);
let updatingCounts = false;
let disposed = false;
const updateLabel = () => {
if (updatingCounts || disposed) return;
formatted.value = `${initialLabel.value}${retryingLabel.value}`;
};
initialLabel.onChanged(updateLabel);
retryingLabel.onChanged(updateLabel);
const updateCounts = () => {
if (disposed) return;
updatingCounts = true;
try {
initialCount.value = source.value.initial;
retryingCount.value = source.value.retrying;
} finally {
updatingCounts = false;
updateLabel();
}
};
source.onChanged(updateCounts);
updateCounts();
return asDisposableReactiveValue(formatted, () => {
if (disposed) return;
disposed = true;
source.offChanged(updateCounts);
initialLabel.offChanged(updateLabel);
retryingLabel.offChanged(updateLabel);
initialLabel.dispose();
retryingLabel.dispose();
});
}
@@ -3,6 +3,7 @@ import { afterEach, beforeEach, describe, expect, it, vi } from "vitest";
import {
STATUS_COUNTER_INACTIVE_LINGER_MS,
createChunkFetchCounterLabel,
createMinimumVisibleActivityCount,
createPaddedCounterLabel,
} from "./StatusBarDisplay.ts";
@@ -137,3 +138,41 @@ describe("createPaddedCounterLabel", () => {
expect(display.value).toBe(" 📄\u20070");
});
});
describe("createChunkFetchCounterLabel", () => {
beforeEach(() => {
vi.useFakeTimers();
});
afterEach(() => {
vi.useRealTimers();
});
it("keeps initial and retry counts separate across an unchanged-total handoff", () => {
const counts = reactiveSource({ initial: 2, retrying: 5 });
const display = createChunkFetchCounterLabel(counts);
const withoutPadding = () => display.value.replace(/\u2007/g, "");
const transitionSnapshots: string[] = [];
const observeTransitions = () => transitionSnapshots.push(withoutPadding());
expect(withoutPadding()).toBe(" 🛄2 🔁5");
display.onChanged(observeTransitions);
counts.value = { initial: 0, retrying: 7 };
expect(withoutPadding()).toBe(" 🛄0 🔁7");
expect(transitionSnapshots).toEqual([" 🛄0 🔁7"]);
display.offChanged(observeTransitions);
vi.advanceTimersByTime(STATUS_COUNTER_INACTIVE_LINGER_MS - 1);
expect(withoutPadding()).toBe(" 🛄0 🔁7");
vi.advanceTimersByTime(1);
expect(withoutPadding()).toBe(" 🔁7");
counts.value = { initial: 0, retrying: 0 };
expect(withoutPadding()).toBe(" 🔁0");
display.dispose();
vi.advanceTimersByTime(STATUS_COUNTER_INACTIVE_LINGER_MS);
counts.value = { initial: 1, retrying: 0 };
expect(withoutPadding()).toBe(" 🔁0");
});
});
+8
View File
@@ -139,6 +139,14 @@ The Harness also measures ID generation with fixed in-memory data on desktop and
The same workflow checks the two remote-activity status boundaries. It first holds a real CouchDB request at the selected fetch implementation and confirms that `🌐N` is visible while `📲` is absent. It then holds the real one-shot replication immediately before its replicator call, confirms that `📲` is visible while no physical request is active, releases it, and requires the finite and bounded activity counts to return to zero, the request and response counts to balance, and both indicators to disappear. Finally, it creates a remote-only chunk, holds the real on-demand fetch immediately before its remote call, makes the same logical active and idle assertions, and verifies that the fetched chunk is written into the local database. These gates make the active states deterministic without replacing the remote request or operation.
`npm run test:e2e:obsidian:focused -- chunk-fetch-retry` checks delayed Chunk availability through a real CouchDB service and Obsidian. It creates a Metadata-only remote fixture, starts ordinary one-shot replication with `readChunksOnline`, and inserts the missing Chunk only after a real fetch has returned an empty result. A pass-through observer records the replicator's call times and results without substituting responses or adding waits. The fixture sets the existing minimum request interval to 500 ms to keep the real retry status observable even if finite completion expedites the final probe. The actual status bar must show zero initial requests (`🛄`) and one retry (`🔁`), and both counts must return to zero after delivery ends. On-demand replication excludes Chunk documents from the ordinary pull, so the delayed Chunk must arrive through the observed fetch and produce the exact Vault content.
If finite replication was already inactive when the initial lookup began, the scenario requires a retry at least two seconds later and no physical request slot occupied during backoff. If finite replication ends during the initial lookup or the following backoff, the retry must instead be a post-completion final probe before the two-second delay would expire. This distinction is determined from the observed finite-count transitions, not an assumed ordering between replication and HTTP completion.
A second case never inserts the Chunk: it requires exactly one retry, a terminal missing notification, released activity, no Vault file, and no further fetch during another retry interval. A third starts a second genuine one-shot replication while the retry remains pending and passively observes its finite count. The final missing lookup must start after that finite operation ends and before the original backoff would expire. No test code changes the finite count. Deterministic Commonlib tests cover longer backoff stages and overlapping-completion races.
Each run uses an isolated Vault and remote database. This is a controlled availability-ordering reproduction, not a reproduction of a particular server's underlying delay or of mobile suspension. When comparing separately built pre-fix and fixed artefacts, retain the exact scenario, package, and bundle revisions: the original implementation ends delivery after the first missing response, while the interim single-retry implementation lacks the split status counts and finite-completion scheduling.
`test:e2e:obsidian:couchdb-manual-setup-workflow` follows the visible onboarding path for the first device when no Setup URI is available. It enters end-to-end encryption and CouchDB details, runs the read-only `Check server requirements` step, requires the prepared fixture to pass without applying a server fix, and lets the onboarding connection test create the named database. After Rebuild completes on the first device, it creates an ordinary note, asks that working device to generate a Setup URI for a second device, completes Fetch there, and verifies a bidirectional note round-trip. The workflow captures each decision point and the expanded server-check result; password controls remain visually masked. It uses an E2EE passphrase beginning with `%`, confirms that the saved settings do not contain it in plain text, and checks that Obsidian restores it after restarting with the first Vault.
The ordinary workflow now checks that all three ID-configuration radio choices are visible, disabled and dimmed while E2EE is off, and fully visible when it is enabled. It also checks that the random key is selected by default for a new Vault, **Keep current configuration** shows its legacy explanation, and the saved key is encrypted locally and transferred by Setup URI. A screenshot of the disabled group is saved as `guide-couchdb-manual-id-generation-disabled.png`. Set `E2E_OBSIDIAN_INDEPENDENT_IDS=true` for the same visible workflow with an explicitly entered, randomly generated source. That variant checks all three nested radio choices, requires a source when no key is saved, retains the saved key when a custom source is empty, rejects an ordinary string in the recovery-code input, restores the same key from a tagged code, verifies that the source is absent from local settings, and checks that both devices compute the same obfuscated document IDs after Setup URI import and Fast Fetch.
@@ -0,0 +1,365 @@
import { mkdir } from "node:fs/promises";
import { join } from "node:path";
import type { Page } from "playwright";
import {
assertCouchDbReachable,
createCouchDbDatabase,
deleteCouchDbDatabase,
loadCouchDbConfig,
makeUniqueDatabaseName,
putCouchDbDocument,
type CouchDbConfig,
} from "../runner/couchdb.ts";
import { discoverObsidianCli, requireObsidianBinary } from "../runner/environment.ts";
import {
assertEqual,
createE2eCouchDbPluginData,
createE2eObsidianDeviceLocalState,
prepareRemote,
waitForLiveSyncCoreReady,
} from "../runner/liveSyncWorkflow.ts";
import { startObsidianLiveSyncSession, type ObsidianLiveSyncSession } from "../runner/session.ts";
import { withObsidianPage } from "../runner/ui.ts";
import { createTemporaryVault } from "../runner/vault.ts";
const observationKey = "__livesyncChunkFetchRetryE2E";
const observationSource = `globalThis[${JSON.stringify(observationKey)}]`;
const retryDelayMs = 2_000;
type FetchAttempt = {
startedAt: number;
finiteTransitionIndex: number;
completedAt?: number;
requestedIds: string[];
returnedIds?: string[];
unavailable?: boolean;
error?: string;
};
type Snapshot = {
attempts: FetchAttempt[];
missingEvents: number[];
replicationDone: boolean;
replicationSucceeded?: boolean;
replicationError?: string;
metadataPresent: boolean;
chunkPresent: boolean;
claimActive: boolean;
currentProcessing: number;
queued: number;
boundedActivity: number;
finiteActivity: number;
finiteTransitions: { at: number; count: number }[];
followupTransitionIndex?: number;
replicationResults: number;
databaseQueue: number;
storageApplying: number;
pendingChunkCount: number;
initialChunkCount: number;
retryChunkCount: number;
statusText: string;
content: string | null;
};
async function snapshot(page: Page): Promise<Snapshot> {
return await page.evaluate(`(async()=>{
const state=${observationSource};
const core=app.plugins.plugins['obsidian-livesync'].core;
const db=core.localDatabase;
const rows=await db.allDocsRaw({keys:[state.metadataId,state.chunkId],include_docs:true});
const present=(id)=>rows.rows.some((row)=>row.id===id&&row.doc&&!row.value?.deleted);
const file=app.vault.getAbstractFileByPath(state.path);
const statusText=document.querySelector('.syncstatusbar')?.textContent??'';
const initialChunkCount=Number(statusText.match(/🛄\\s*(\\d+)/u)?.[1]??0);
const retryChunkCount=Number(statusText.match(/🔁\\s*(\\d+)/u)?.[1]??0);
return {
attempts:state.attempts,missingEvents:state.missingEvents,
replicationDone:state.replicationDone,replicationSucceeded:state.replicationSucceeded,
replicationError:state.replicationError,metadataPresent:present(state.metadataId),
chunkPresent:present(state.chunkId),
claimActive:db.managers.chunkManager.deliveryCoordinator.isClaimActiveFor(state.chunkId),
currentProcessing:db.managers.chunkFetcher.currentProcessing,
queued:db.managers.chunkFetcher.queue.length,
boundedActivity:core.services.replicator.boundedRemoteActivityCount.value,
finiteActivity:core.services.replicator.finiteReplicationActivityCount.value,
finiteTransitions:state.finiteTransitions,followupTransitionIndex:state.followupTransitionIndex,
replicationResults:core.services.replication.replicationResultCount.value,
databaseQueue:core.services.replication.databaseQueueCount.value,
storageApplying:core.services.replication.storageApplyingCount.value,
initialChunkCount,retryChunkCount,pendingChunkCount:initialChunkCount+retryChunkCount,statusText,
content:file?await app.vault.read(file):null,
};
})()`);
}
async function waitForSnapshot(
page: Page,
predicate: (state: Snapshot) => boolean,
stage: string,
timeoutMs = 20_000
): Promise<Snapshot> {
const deadline = Date.now() + timeoutMs;
let state: Snapshot;
do {
state = await snapshot(page);
if (predicate(state)) return state;
if (state.replicationError || state.attempts.some((attempt) => attempt.error || attempt.unavailable)) {
throw new Error(`The real CouchDB operation failed during ${stage}: ${JSON.stringify(state)}`);
}
await new Promise((resolve) => setTimeout(resolve, 50));
} while (Date.now() < deadline);
throw new Error(`Timed out during ${stage}: ${JSON.stringify(state)}`);
}
function isIdle(state: Snapshot): boolean {
return (
state.replicationDone &&
!state.claimActive &&
state.currentProcessing === 0 &&
state.queued === 0 &&
state.boundedActivity === 0 &&
state.finiteActivity === 0 &&
state.replicationResults === 0 &&
state.databaseQueue === 0 &&
state.storageApplying === 0 &&
state.pendingChunkCount === 0
);
}
async function runScenario(
page: Page,
couchDb: CouchDbConfig,
dbName: string,
label: "delayed-arrival" | "permanently-missing" | "finite-completion"
): Promise<void> {
const delayedArrival = label === "delayed-arrival";
const expediteFinalProbe = label === "finite-completion";
const path = `chunk-fetch-${label}.md`;
const chunkId = `h:e2e-chunk-fetch-${label}`;
const content = `# Chunk fetch retry\n${label}\n`;
const metadataId = await page.evaluate<string>(
`app.plugins.plugins['obsidian-livesync'].core.services.path.path2id(${JSON.stringify(path)})`
);
const now = Date.now();
// Only Metadata is initially present. The Chunk is a separate real CouchDB document.
await putCouchDbDocument(couchDb, dbName, {
_id: metadataId,
path,
type: "plain",
ctime: now,
mtime: now,
size: Buffer.byteLength(content),
children: [chunkId],
eden: {},
});
await page.evaluate(`(()=>{
if(${observationSource}) throw new Error('A Chunk fetch observer is already installed.');
const core=app.plugins.plugins['obsidian-livesync'].core;
const settings=core.services.setting.currentSettings();
if(!settings.readChunksOnline||settings.useOnlyLocalChunk||settings.liveSync||settings.periodicReplication) {
throw new Error('The fixture requires one-shot replication with on-demand Chunk reads.');
}
const replicator=core.services.replicator.getActiveReplicator();
const original=replicator.fetchRemoteChunks;
if(typeof original!=='function') throw new Error('The real replicator has no Chunk fetch method.');
const manager=core.localDatabase.managers.chunkManager;
const observer=new AbortController();
const finiteCount=core.services.replicator.finiteReplicationActivityCount;
const state={path:${JSON.stringify(path)},metadataId:${JSON.stringify(metadataId)},
chunkId:${JSON.stringify(chunkId)},attempts:[],missingEvents:[],replicationDone:false,
finiteTransitions:[{at:Date.now(),count:finiteCount.value}]};
const observeFinite=()=>state.finiteTransitions.push({at:Date.now(),count:finiteCount.value});
finiteCount.onChanged(observeFinite);
${observationSource}=state;
state.restore=()=>{replicator.fetchRemoteChunks=original;observer.abort();finiteCount.offChanged(observeFinite);};
manager.addListener('missingChunkRemote',(id)=>{
if(id===state.chunkId) state.missingEvents.push(Date.now());
},{signal:observer.signal});
// Observe the real HTTP-backed method without changing its result or adding a wait.
replicator.fetchRemoteChunks=async function(...args){
if(!args[0].includes(state.chunkId)) return await original.apply(this,args);
const attempt={startedAt:Date.now(),finiteTransitionIndex:state.finiteTransitions.length,requestedIds:[...args[0]]};
state.attempts.push(attempt);
try{
const result=await original.apply(this,args);
attempt.unavailable=result===false;
attempt.returnedIds=Array.isArray(result)?result.map((chunk)=>chunk._id):[];
return result;
}catch(error){
attempt.error=String(error);
throw error;
}finally{
attempt.completedAt=Date.now();
}
};
state.startReplication=async()=>{
state.replicationDone=false;
try{state.replicationSucceeded=!!(await core.services.replication.replicate(true));}
catch(error){state.replicationError=String(error);}
finally{state.replicationDone=true;}
};
state.replication=state.startReplication();
})()`);
try {
const first = await waitForSnapshot(
page,
(state) => !!state.attempts[0]?.completedAt,
"initial missing response"
);
assertEqual(first.metadataPresent, true, "The Metadata did not arrive through real replication.");
assertEqual(first.chunkPresent, false, "The Chunk was already available locally before its delayed arrival.");
assertEqual(
first.attempts[0].unavailable,
false,
"The first fetch failed instead of returning a missing Chunk."
);
assertEqual(first.attempts[0].returnedIds?.length, 0, "The first fetch did not reproduce a missing Chunk.");
if (delayedArrival) {
await putCouchDbDocument(couchDb, dbName, { _id: chunkId, type: "leaf", data: content });
}
console.log(`${label}: first real response ${JSON.stringify(first)}`);
assertEqual(
first.missingEvents.length,
0,
"The first missing response ended delivery before the delayed retry."
);
assertEqual(first.claimActive, true, "The first missing response released the delivery claim.");
assertEqual(first.content, null, "The file was materialised before its missing Chunk arrived.");
const retryWaiting = await waitForSnapshot(
page,
(state) => state.initialChunkCount === 0 && state.retryChunkCount === 1,
"separate retry status during the retry delay",
1_000
);
assertEqual(retryWaiting.attempts.length, 1, "The pending count appeared only after the retry started.");
const completedDuringInitialLookup = retryWaiting.finiteTransitions
.slice(first.attempts[0].finiteTransitionIndex)
.some((transition) => transition.count === 0);
if (!completedDuringInitialLookup) {
assertEqual(retryWaiting.currentProcessing, 0, "Backoff retained a physical request slot.");
}
console.log(`${label}: retry waiting ${JSON.stringify(retryWaiting)}`);
if (expediteFinalProbe) {
await waitForSnapshot(
page,
(state) => state.replicationDone && state.finiteActivity === 0 && state.attempts.length === 1,
"first finite replication completion before the scheduled retry",
1_000
);
// A second genuine one-shot replication ends during backoff; counts are observed, never synthesised.
await page.evaluate(`(()=>{
const state=${observationSource};
state.followupTransitionIndex=state.finiteTransitions.length;
state.replication=state.startReplication();
})()`);
}
const completed = await waitForSnapshot(page, isIdle, "delivery and reflection quiescence");
assertEqual(completed.replicationSucceeded, true, "One-shot replication did not complete successfully.");
assertEqual(completed.attempts.length, 2, "The missing Chunk must be fetched exactly twice.");
const [initial, retry] = completed.attempts;
const retryDelay = retry.startedAt - initial.completedAt!;
if (expediteFinalProbe) {
const transitions = completed.finiteTransitions.slice(completed.followupTransitionIndex);
const ended = transitions.find(
(transition, index) => index > 0 && transition.count === 0 && transitions[index - 1].count > 0
);
if (!ended)
throw new Error(`The second real finite replication was not observed: ${JSON.stringify(transitions)}`);
if (retry.startedAt < ended.at) throw new Error("The final probe began before finite replication ended.");
if (retryDelay >= retryDelayMs)
throw new Error(`Finite completion did not interrupt backoff: ${retryDelay} ms.`);
} else {
const completion = completed.finiteTransitions
.slice(initial.finiteTransitionIndex)
.find((transition) => transition.count === 0);
if (completion) {
if (retry.startedAt < completion.at) throw new Error("The final probe preceded finite completion.");
if (retryDelay >= retryDelayMs) {
throw new Error(`Finite completion did not expedite the initial missing result: ${retryDelay} ms.`);
}
} else if (retryDelay < retryDelayMs) {
throw new Error(`The retry started too early: ${retryDelay} ms.`);
}
}
assertEqual(retry.requestedIds.join(","), chunkId, "The retry requested an unexpected Chunk.");
assertEqual(retry.unavailable, false, "The retry failed to contact the real remote.");
assertEqual(retry.returnedIds?.join(","), delayedArrival ? chunkId : "", "The retry returned unexpected data.");
assertEqual(
completed.missingEvents.length,
delayedArrival ? 0 : 1,
"Unexpected terminal missing notifications."
);
assertEqual(completed.chunkPresent, delayedArrival, "The local Chunk persistence result was unexpected.");
assertEqual(completed.content, delayedArrival ? content : null, "The Vault file content was unexpected.");
if (!delayedArrival) {
// Observe one more retry interval after quiescence to reject an unbounded retry loop.
await new Promise((resolve) => setTimeout(resolve, retryDelayMs + 100));
const settled = await snapshot(page);
assertEqual(isIdle(settled), true, "The permanently missing delivery became active again.");
assertEqual(settled.attempts.length, 2, "The permanently missing Chunk was retried again.");
}
console.log(`${label}: passed; retry after ${retryDelay} ms; ${JSON.stringify(completed)}`);
} catch (error) {
console.error(`${label}: ${JSON.stringify(await snapshot(page))}`);
const diagnostics = process.env.E2E_OBSIDIAN_DIAGNOSTICS_DIR ?? "/tmp/obsidian-livesync-e2e";
await mkdir(diagnostics, { recursive: true });
await page.screenshot({ path: join(diagnostics, `chunk-fetch-${label}.failure.png`), fullPage: true });
throw error;
} finally {
await page.evaluate(`(()=>{${observationSource}?.restore();delete ${observationSource};})()`);
}
}
async function main(): Promise<void> {
const binary = requireObsidianBinary();
const cli = discoverObsidianCli();
if (!cli.binary) throw new Error(`Could not find obsidian-cli. Checked: ${cli.checked.join(", ")}`);
const couchDb = await loadCouchDbConfig();
await assertCouchDbReachable(couchDb);
const dbName = makeUniqueDatabaseName(couchDb.dbPrefix, "chunk-fetch-retry");
const vault = await createTemporaryVault("obsidian-livesync-chunk-fetch-");
let session: ObsidianLiveSyncSession | undefined;
try {
await createCouchDbDatabase(couchDb, dbName);
session = await startObsidianLiveSyncSession({
binary,
cliBinary: cli.binary,
vault,
pluginData: createE2eCouchDbPluginData(
{ ...couchDb, dbName },
{
encrypt: false,
usePathObfuscation: false,
showStatusOnStatusbar: true,
// Keep the real retry status observable even when finite completion expedites the final probe.
minimumIntervalOfReadChunksOnline: 500,
periodicReplication: false,
syncOnFileOpen: false,
syncOnEditorSave: false,
syncAfterMerge: false,
}
),
localStorageEntries: createE2eObsidianDeviceLocalState(vault.name),
});
await waitForLiveSyncCoreReady(cli.binary, session.cliEnv);
await prepareRemote(cli.binary, session.cliEnv);
await withObsidianPage(session.remoteDebuggingPort, async (page) => {
await runScenario(page, couchDb, dbName, "delayed-arrival");
await runScenario(page, couchDb, dbName, "permanently-missing");
await runScenario(page, couchDb, dbName, "finite-completion");
});
} finally {
if (session) await session.app.stop();
await vault.dispose();
await deleteCouchDbDatabase(couchDb, dbName);
}
}
main().catch((error: unknown) => {
console.error(error instanceof Error ? error.stack : error);
process.exitCode = 1;
});
+1
View File
@@ -17,6 +17,7 @@ const focusedScenarios = new Set([
"p2p-pane",
"vault-reflection",
"couchdb-upload",
"chunk-fetch-retry",
"couchdb-manual-setup-workflow",
"cli-to-obsidian-sync",
"minio-upload",
+3
View File
@@ -33,6 +33,9 @@ Earlier releases remain available in the 1.0 release history, the 1.0 preview hi
- We can now keep using an E2EE passphrase beginning with `%` after restarting Obsidian. (#1221)
- LiveSync encrypts it before saving the settings. If an earlier version saved it in plain text, re-enter the passphrase used to encrypt the existing data after updating. Treat that passphrase as exposed if the affected `data.json` was shared.
- A receiving device now retries an unavailable CouchDB Chunk when file Metadata arrives before that Chunk is visible, helping rapid edits reach the Vault after an initial on-demand lookup misses it. (#1224)
- Retries start after two seconds and continue with increasing delays while finite replication is active. When it ends, LiveSync checks locally and makes a final lookup if needed, without waiting out the remaining retry delay.
- We can now distinguish initial on-demand Chunk requests (`🛄`) from retries (`🔁`) in the status bar. These replace `🧩`; each pending Chunk appears in one category, including while a retry is waiting.
### Synchronisation and storage