We stored user configuration in a managed document store — flexible JSON objects, schema that could grow as the product changed. One field was a map: opaque string keys to small records. Saves came in through an API that often needed to add just one key.
The naive path is read-modify-write: fetch the map, set your key, write the whole field back. Two of those at once can drop a key. Each request starts from a snapshot that is already stale by the time it writes.
Don’t PUT the object
A full object replace is worse. The document also held sibling fields — things like defaultFolderId and schemaVersion. Putting the whole object from a partial client view can wipe those defaults. The store exposed JSON Patch for a reason:
PATCH /objects/{id}
Content-Type: application/json-patch+json
Patch only the field you mean to change. Leave everything else alone.
Compare-and-swap with test + replace
RFC 6902 defines a test operation: if the current value at a path does not match, the whole patch is rejected and nothing changes. Pair that with replace and you get compare-and-swap without a separate version field. The client does not send a version number; the value is the version.
On POST /v1/saved-skills, after validating the new entry, and assuming the config object already exists for the user:
- Read
user_config_{userId}and keepskillMapexactly as stored. - Build a new map by setting that one key — entry file id and
createdAt— and leave every other key as it was. - Patch only
skillMap, guarded by the value just read.
[
{
"op": "test",
"path": "/skillMap",
"value": {
"111": { "fileId": "9001", "createdAt": "2026-09-01T12:00:00Z" }
}
},
{
"op": "replace",
"path": "/skillMap",
"value": {
"111": { "fileId": "9001", "createdAt": "2026-09-01T12:00:00Z" },
"222": { "fileId": "9002", "createdAt": "2026-09-22T15:00:00Z" }
}
}
]
The test value must be the map returned by the read. The replace value is that map with this request’s one key applied. If skillMap is missing, use add for /skillMap and skip test.
Two concurrent saves
The client already has a skillMap. It fires two requests at once, each appending a different key. Those calls hit our microservice — skipping GraphQL, load balancers, and other layers in between for this explanation. The microservice talks to a document service that manages Bigtable for us.
sequenceDiagram
autonumber
participant Client
participant A as Request A
participant B as Request B
participant Doc as DocumentService
Note over Client,Doc: Existing skillMap = S
par Save skill A
Client->>A: POST /v1/saved-skills (A)
A->>Doc: Read skillMap
Doc-->>A: Snapshot S
and Save skill B
Client->>B: POST /v1/saved-skills (B)
B->>Doc: Read skillMap
Doc-->>B: Snapshot S
end
Note over Client,Doc: Both requests based on snapshot S
A->>Doc: PATCH test(S) replace(S+A)
Doc-->>A: 200 OK
A-->>Client: 200 OK
B->>Doc: PATCH test(S) replace(S+B)
Doc-->>B: 409 Conflict
alt Retry on conflict
B->>Doc: Read skillMap
Doc-->>B: Snapshot S+A
B->>Doc: PATCH test(S+A) replace(S+A+B)
Doc-->>B: 200 OK
B-->>Client: 200 OK
else No retry
B-->>Client: 409 Conflict
end
Without test, both requests can read the same snapshot, each build a map with only their own new key, and both patch /skillMap. The later write wins; the other key is gone. Neither call knows the map changed under it.
With test, the first request’s test plus replace succeeds and the client gets success. The second request’s test still asserts the old map; the document service rejects the patch with 409. Nothing is written for that attempt.
Retry on conflict
If test fails, the document service returns 409 — another write landed first. Re-read and run the same decision again: build the new map from the fresh snapshot, then test plus replace. Do not write the stale map. A few attempts is enough. If they are exhausted, return a retryable 503. The caller’s retry is safe when a save that already landed is treated as an already-saved conflict.
What does not collide
A create race stays separate: a 409 on create means the instance already exists, then fall into this patch path. An update that patches only defaultFolderId does not collide with this write — different paths, different fields.
The lesson I keep: when the store speaks RFC 6902, use test as your lock. Read-modify-write of a whole map is how concurrent saves quietly lose each other’s work.