Skip to content
indraft
Start free

Cookbook

Keep a spreadsheet or a warehouse in step

Somebody always needs a copy: the finance sheet, the warehouse, the dashboard the board looks at. This is the read that makes an outside copy possible without it drifting.

The situation

Somebody outside the CRM needs the data. The naive version pulls everything every night, which works until it does not: it gets slower every month, it cannot tell you what changed, and if it fails on Tuesday you either re-pull the world or quietly lose a day.

The version that keeps working asks a different question. Not "what is there" but "what has happened since the last time I asked".

What must be true when you are done

  • Each run asks only for what changed since the previous run.
  • A run that fails can be repeated without skipping or duplicating anything.
  • Each change says who made it and how, not just what the value now is.
  • You know how far back you can ask, and what happens if you fall behind.

This is an automation, not an agent

Worth stating before the mechanics. A nightly copy should be identical every night, run when nobody is watching, and never exercise judgment. That is the definition of the work an agent is the wrong tool for, and deciding which is which is worth reading if you are choosing. Everything below is configuration in whatever automation tool you already run.

Ask what changed

One read returns the changes and a marker. The marker is the whole mechanism: you keep it, and hand it back next time.

Agent

changes create_task task created via rest title Escalation: data load stalled twice actor the credential your job runs as at 25 Aug 00:38next marker aWRmYzE6bXV0XzAxTTBWNUpB...oldest kept 25 Aug 00:18

Three things in that response earn their place. Each entry names the operation and the outcome, so a change that was refused does not look like one that applied. Each entry names the actor and the route it came in by, so a row your automation wrote is distinguishable from one a person typed, which is what stops a two-way setup echoing itself forever. And the marker is opaque on purpose: you store it, you do not parse it.

Store the marker, not the timestamp

This is where these jobs are usually got wrong. The instinct is to record the time of the last run and ask for everything after it, which loses any change that landed during the run and duplicates any that landed in the same second. The marker has no such gap: it names a position rather than a moment.

In practice that is three steps in your automation tool. Read the marker you saved. Ask for changes since it. Save the new marker only after the rows are safely written. If the write fails, the marker is untouched and the next run asks for the same changes again, which is the behaviour you want.

Know what happens if you fall behind

The response tells you the oldest change still kept. That is the number to pay attention to, because it is the answer to "what if my job is broken for a fortnight". If your marker is older than the oldest retained change, an incremental catch-up cannot be correct, and the honest recovery is a full reload rather than resuming and hoping.

Have your automation compare the two and alert rather than silently continuing. A sync that quietly resumes past a gap is worse than one that stops, because the gap is invisible in the destination and shows up as a number nobody can reconcile.

Taking the whole thing, rather than the changes

There is also a full export: every record, your own object types, and the history, as one file. It is deliberately harder to reach than the change feed. A credential can only run it if it was explicitly created with permission to export AND the person who created it is an owner, so it is not something a token left in a configuration file can do quietly. It is also bounded: above 100 MB in one file it refuses rather than returning part of your workspace, and we run it for you in that case. That bound is another reason a recurring copy belongs on the change feed.

That is the right trade for the operation that takes everything at once, and it is why the recurring job above is built on the change feed instead. Use the export for the deliberate acts, a migration or a backup you have decided to take, and the feed for anything that runs on a schedule.

Verify it

  • Run it twice with no changes in between. The second run should return nothing and write nothing. If it re-writes rows, you are storing a timestamp rather than the marker.
  • Kill it mid-run on purpose. Then run it again. Nothing should be missing and nothing duplicated. This is the only test that tells you whether the marker is saved at the right moment.
  • Change one record and watch one row move. Not a full reload; one row. If a single edit produces a large batch, you are pulling rather than following.
  • Check the oldest-kept alert fires. Set the marker to something ancient and confirm your job complains instead of proceeding.

Run it again

On whatever schedule the destination needs, which is usually far less often than people set it to. If the copy is feeding something people watch during the day, having Indraft tell your automation the moment something happens is the better shape than polling more frequently.