Skip to content

feat: named child runs that survive migration and resurrection - #759

Open
barjin wants to merge 24 commits into
masterfrom
feat/track-child-runs
Open

barjin wants to merge 24 commits into
masterfrom
feat/track-child-runs

Conversation

@barjin

@barjin barjin commented Sep 30, 2026

Copy link
Copy Markdown
Member

Adds a runName option to Actor.start, call and callTask that reattaches to the recorded child run after a restart, and Actor.childRuns() to read them.

Closes #738, closes #740

@github-actions

github-actions Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

⚠️ There are broken links in the documentation.

See more at https://github.com/apify/apify-sdk-js/actions/runs/37911125694#summary-113756252850

@barjin barjin left a comment

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A few ideas / comments where this implementation differs from the Orchestrator ⬇️

Comment thread src/actor.ts
Comment thread src/child_run_tracker.ts
Comment thread src/actor.ts Outdated
const tracked = await this.#childRunTracker.get(runName);
const trackedRun = tracked ? await client.run(tracked.runId).get() : undefined;

if (trackedRun && ['SUCCEEDED', 'READY', 'RUNNING'].includes(trackedRun.status)) {

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Every other run status (aborting / timed-out) is currently replaced by a completely new run instance.

Should we attempt to resurrect these?

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Bonus question for resurrected runs - shall we try implementing the 'inherit' timeout mode (pass the remaining timeout to the resurrected child run)?

There's no Actor.resurrect() method in SDK now (is this intentional?)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Every other run status (aborting / timed-out) is currently replaced by a completely new run instance. Should we attempt to resurrect these?

Yes, I'd rather resurrect them.

Bonus question for resurrected runs - shall we try implementing the 'inherit' timeout mode (pass the remaining timeout to the resurrected child run)?

It makes sense. Let's try to implement it.

There's no Actor.resurrect() method in SDK now (is this intentional?)

I'm not aware of this being intentionally missing, so I'd add it and use it (Python SDK is missing it as well).

@barjin barjin self-assigned this Sep 30, 2026
@barjin
barjin marked this pull request as ready for review September 30, 2026 09:55
@barjin
barjin requested review from B4nan and janbuchar September 30, 2026 09:55
@apify-service-account apify-service-account added tested Temporary label used only programatically for some analytics. t-tooling Issues with this label are in the ownership of the tooling team. labels Sep 30, 2026
Comment thread src/actor.ts Outdated
Comment thread src/child_run_tracker.ts
Comment thread src/child_run_tracker.ts
Comment thread src/child_run_tracker.ts Outdated
Comment thread src/actor.ts
barjin and others added 4 commits October 6, 2026 13:35
Co-authored-by: Martin Adámek <banan23@gmail.com>
Co-authored-by: Martin Adámek <banan23@gmail.com>
Co-authored-by: Martin Adámek <banan23@gmail.com>
`Date` and `URL` inputs were hashed as `{}` and function inputs were dropped, so inputs differing only in those
values got the same checksum and the earlier run was resumed. Hash the JSON form `apify-client` sends instead,
and sort keys by code units, as `localeCompare` depends on the locale.
@barjin
barjin requested a review from B4nan October 6, 2026 12:56
@B4nan

B4nan commented Oct 6, 2026

Copy link
Copy Markdown
Member

Lets compare this with apify/apify-sdk-python#1149

@B4nan

B4nan commented Oct 6, 2026 •

Copy link
Copy Markdown
Member

Compared with the Python counterpart, which is split into apify/apify-sdk-python#1149 (named runs) and apify/apify-sdk-python#1151 (Actor.child_runs()):

JS (this PR) Python
Option runName name
Methods start, call, callTask start, call (no call_task)
KVS key CHILD_RUNS APIFY_CHILD_RUNS
Stored record runId, status, startedAt, checksum, history[{runId, status, startedAt}] actorId, runId, previousRunIds
Name reused for a different request throws if the Actor/task or input differs ValueError only if the actor ID differs, a new input is accepted
ABORTED / TIMED-OUT replaced by a new run resurrected (with build, memory, timeout and max charge passed on)
ABORTING / TIMING-OUT a new run starts right away, while the old one is still finishing waits for it to finish, then resurrects it
FAILED / run gone new run, old one kept in history same
Concurrent calls with the same name per-name promise chain per-name asyncio.Lock
call log on a resumed run fromStart: !resumed same, plus a status message watcher; a SUCCEEDED run returns without streaming
childRuns() stored snapshot with the last observed status, no API calls fetches every run from the API; run: None if it's gone
Malformed stored data not checked raises ValueError
Tests unit unit + e2e

Things we should settle across both SDKs:

  1. KVS key and record format. These should match, so the stored data looks the same whichever SDK wrote it.
  2. Resurrect or replace aborted/timed-out runs (your question above). Either way, we should wait out ABORTING / TIMING-OUT before starting a replacement, otherwise the old and new runs overlap.
  3. Whether the name check compares input (JS) or just the Actor (Python).
  4. Whether childRuns() fetches fresh status from the API.
  5. Python has no named call_task yet.

@vdusek

vdusek commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

Python PRs (apify/apify-sdk-python#1149, apify/apify-sdk-python#1151) have been changed to match JS behavior in two places:

  • the option is now named run_name,
  • and call_task supports it as well.

I'm still thinking about the rest.

Comment thread src/child_run_tracker.ts Outdated
Comment thread src/actor.ts Outdated
Comment thread src/child_run_tracker.ts
Comment thread src/child_run_tracker.ts

@janbuchar janbuchar left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I have a bunch of code cleanup suggestions, but nothing critical. The PR looks fine to me.

Comment thread test/apify/actor.test.ts Outdated
Comment thread test/apify/actor.test.ts Outdated
Comment thread test/apify/actor.test.ts Outdated
Comment thread test/apify/actor.test.ts Outdated
Comment thread test/apify/actor.test.ts Outdated
Comment thread src/child_run_tracker.ts Outdated
Comment thread src/actor.ts
Comment thread src/actor.ts Outdated
Comment thread src/actor.ts Outdated
@barjin
barjin force-pushed the feat/track-child-runs branch from bc1cc0b to c83bcc6 Compare October 9, 2026 09:01

@vdusek vdusek left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

See this separate comment #759 (comment).

Comment thread src/child_run_tracker.ts

return this.locks.get(runName)!.runExclusive(async () => {
const tracked = await this.get(runName);
const trackedRun = tracked ? await client.run(tracked.runId).get() : undefined;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A 404 here right after a start (e.g. a concurrent call, or a migration moments after start()) can just be a lagging API replica, and we would start a duplicate run. Python retries get() for 3 s before treating the run as LOST.

See apify/apify-sdk-python#1149 (comment).

✍️ Drafted by Claude Code

Comment thread src/child_run_tracker.ts
const tracked = await this.get(runName);
const trackedRun = tracked ? await client.run(tracked.runId).get() : undefined;

if (trackedRun && ['SUCCEEDED', 'READY', 'RUNNING'].includes(trackedRun.status)) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

An ABORTING / TIMING-OUT run falls through to start(), so the old and new runs overlap for a while. Python waits for such a run to finish (waitForFinish()) and then handles it as ABORTED / TIMED-OUT.

✍️ Drafted by Claude Code

Comment thread src/child_run_tracker.ts
const trackedRun = tracked ? await client.run(tracked.runId).get() : undefined;

if (trackedRun && ['SUCCEEDED', 'READY', 'RUNNING'].includes(trackedRun.status)) {
await this.verifyRequest(runName, request);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The checksum is only checked on this reattach path, so a name whose recorded run failed or got lost accepts a different input, and the new checksum overwrites the old one. Python raises in that case too, so a name stays bound to one Actor/task + input. Which way we want to go?

✍️ Drafted by Claude Code

Comment thread src/child_run_tracker.ts
return { run: trackedRun, resumed: true };
}

const run = await start();

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ABORTED / TIMED-OUT runs end up here and get replaced by a new run, but we agreed to resurrect them. Python resurrects them with the call's build, memory, timeout and max charge. JS needs an Actor.resurrect() for that first.

This can be addressed once this #769 is merged.

✍️ Drafted by Claude Code

@vdusek

vdusek commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

Compared with the Python counterpart (apify/apify-sdk-python#1149 for named runs, apify/apify-sdk-python#1151 for child_runs). Python now matches this PR on the option name, supported methods, KVS key, record format, checksum, per-name locking, and the childRuns shape (sync, run clients keyed by name, custom-token client kept in-process).

Full comparison table:

Area JS (this PR) Python Aligned Notes
Option name runName run_name ✅ Python renamed name → run_name
Supported methods start, call, callTask start, call, start_task, call_task ❌ JS has no Actor.startTask() yet, see #764
Record format runId, status, startedAt, checksum, history[] same ✅
Checksum sha256 of sorted-key JSON of {type, id, input} same, byte-equal ✅
Write timing right after start right after start ✅
READY / RUNNING reattach reattach ✅
SUCCEEDED return as is return as is ✅
FAILED / run gone new run, old one moved to history same ✅
ABORTED / TIMED-OUT new run resurrected with the call's build, memory, timeout and max charge ❌ Resurrecting is cheaper and was agreed, see r4230501550
ABORTING / TIMING-OUT new run right away, overlapping the old one waits for it to finish, then resurrects ❌ Avoids two children running under one name, see r4230501533
Actor.resurrect() missing added in apify/apify-sdk-python#1165 ❌ Needed for the ABORTED / TIMED-OUT row
404 on the recorded run LOST immediately retried for 3 s first ❌ A lagging API replica can otherwise start a duplicate, see r4230501519
Different input, run reattached throws raises ✅
Different input, run replaced allowed, checksum overwritten raises ❌ Open question whether a name stays bound to one input, see r4230501542
Concurrent calls under one name per-name lock per-name lock ✅
childRuns API sync getter, Record<string, RunClient> sync property, dict[str, RunClientAsync] ✅ Python aligned in apify/apify-sdk-python#1151
Registry load in init() in __aenter__ ✅
Custom token in childRuns kept in-process, default client after migration same ✅ ported from Python
History in childRuns KVS record only KVS record only ✅
Access before init warns, returns {} raises ✅ each SDK's existing convention, fine as is
Malformed record not validated ValueError at init ❌ minor
Max charge / max_items forwarded to the start forwarded to the start and to the resurrection ✅
Tests unit unit + e2e (reboot → reattach) ❌

✍️ Drafted by Claude Code

vdusek added a commit to apify/apify-sdk-python that referenced this pull request Oct 9, 2026
`Actor.start`, `Actor.call`, `Actor.start_task` and `Actor.call_task`
accept an optional `run_name`. Right after the child starts, the SDK
records it in the default key-value store under `__ACTOR_CHILD_RUNS`.
After a migration or resurrection of the parent, the same call finds the
recorded run and reuses it:

- `READY` / `RUNNING`: reattaches to it
- `SUCCEEDED`: returns it as is
- `ABORTED` / `TIMED-OUT`: resurrects it with the call's build, memory,
timeout, `max_items` and max charge (an `ABORTING` / `TIMING-OUT` run is
waited for first)
- `FAILED`, or the run no longer exists: starts a new run and moves the
old one to `history`

The record has the same shape as in the JS SDK (apify/apify-sdk-js#759):
`runId`, `status`, `startedAt`, `checksum` and `history`. The checksum
covers the Actor or task ID and the input, computed the same way as in
JS. Reusing a name for a different Actor, task or input raises
`ValueError`. Concurrent calls under one name in the same process share
a single run.

A named `call` on a reattached or resurrected run streams only new log
lines. If the API reports a recorded run as missing, the SDK retries for
3 seconds before replacing it, because a run started moments ago may not
be on every API replica yet.

A hard kill between the platform starting the child and the registry
write can still orphan the child. Closing that gap needs an idempotency
key on the run-start endpoint.

Closes: #1127

*✍️ Drafted by Claude Code*
vdusek added a commit to apify/apify-sdk-python that referenced this pull request Oct 9, 2026
`Actor.child_runs` is a sync property that maps each run name to a
`RunClientAsync` for the current run under that name. It has the same
shape as `Actor.childRuns` in the JS SDK (apify/apify-sdk-js#759). The
client covers the run and its storages:

```python
run = await Actor.child_runs['scrape-eu'].wait_for_finish()
```

The SDK reads the registry from the default key-value store when the
Actor initializes, so runs recorded before a migration or resurrection
are included and reading the property makes no API calls. That adds one
record read to every init. A malformed record fails the init and leaves
the Actor uninitialized. Runs started without a `run_name` aren't
tracked. Runs that were replaced under a name stay in the `history` of
the `__ACTOR_CHILD_RUNS` record and aren't exposed.

A run started or reattached with a custom `token` gets a client with
that token. A run recorded before a migration and not reattached since
gets the default client, because the token isn't stored.

It also fixes a duplicate `max_items` deprecation warning from #1149: a
named start that resurrects its recorded run no longer warns a second
time from inside the SDK.

Docs are in #1160.

Closes: #1129

*✍️ Drafted by Claude Code*
Comment thread src/actor.ts
Comment on lines +892 to +904
const { waitSecs, log, ...startOptions } = rest;
const { run, resumed } = await this.#childRunTracker.startOrResumeChildRun(
client,
runName,
{ type: 'actor', id: actorId, input },
async () => client.actor(actorId).start(input, { ...startOptions, runTimeoutSecs }),
);

// The earlier part of a resumed run's log was already redirected before the migration.
const streamedLog = await client.run(run.id).getStreamedLog({ toLog: log, fromStart: !resumed });
streamedLog?.start();
try {
const finishedRun = await client.run(run.id).waitForFinish({ waitSecs });

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[medium] With runName, signal and timeoutSecs only reach start(). The unnamed path (client call()) also passes them to waitForFinish / getStreamedLog and defaults timeoutSecs: 'noTimeout'. Aborting the signal doesn't stop Actor.call(..., { runName, signal }); it hangs until the child finishes. And start() falls back to the 'medium' request timeout, which can cut off a start with waitForResources. Same in callTask() below. Fix: forward { waitSecs, timeoutSecs, signal } to waitForFinish, signal to getStreamedLog, and default timeoutSecs like the client does.

Repro (with runName times out, without passes)
import { createServer } from 'node:http';
import type { AddressInfo } from 'node:net';
import { MemoryStorageBackend } from '@crawlee/core';
import { Configuration } from 'apify';
import type { Run as ActorRun } from 'apify-client';
import { ActorClient } from 'apify-client';
import { createIsolatedActor } from '../createIsolatedActor.js';

const server = createServer(() => {});
let apiBaseUrl: string;
beforeAll(async () => {
    await new Promise<void>((resolve) => server.listen(0, '127.0.0.1', resolve));
    apiBaseUrl = `http://127.0.0.1:${(server.address() as AddressInfo).port}/`;
});
afterAll(() => { server.closeAllConnections(); server.close(); });

const run = { id: 'child-run', status: 'RUNNING', startedAt: new Date() } as ActorRun;

test.each([['without runName', {}], ['with runName', { runName: 'child' }]])(
    'call() %s rejects once its abort signal fires',
    async (_label, extra) => {
        vitest.spyOn(ActorClient.prototype, 'start').mockResolvedValue(run);
        const { actor } = createIsolatedActor({
            config: new Configuration({ apiBaseUrl, token: 'token' }),
            storageClient: new MemoryStorageBackend(),
        });
        const signal = AbortSignal.timeout(200);
        await expect(actor.call('some-actor', {}, { signal, log: null, ...extra })).rejects.toThrow();
    },
    3000,
);

Comment thread src/actor.ts
Comment on lines +1012 to +1019
const { waitSecs, ...startOptions } = rest;
const { run } = await this.#childRunTracker.startOrResumeChildRun(
client,
runName,
{ type: 'task', id: taskId, input },
async () => client.task(taskId).start(input, { ...startOptions, runTimeoutSecs }),
);
const finishedRun = await client.run(run.id).waitForFinish({ waitSecs });

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[medium] Same as in call(): signal / timeoutSecs never reach waitForFinish.

Comment thread src/actor.ts
});
log.debug(`Default storages purged`);

await this.#childRunTracker.load();

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[low] This adds a getRecord('__ACTOR_CHILD_RUNS') (plus a store get) to every Actor.init() on the platform, also for Actors that never use runName, and init fails if the read fails. Is that cost intended for the sync childRuns getter? Alternative: load lazily on the first named start/call/callTask.

Comment thread src/actor.ts
*
* If a `READY`, `RUNNING` or `SUCCEEDED` run is already tracked under this name, it is returned instead of
* starting a new one, including after a migration or resurrection of this run. A run that failed, was aborted
* or timed out is replaced by a new one. Reusing the name for a different Actor, task or input throws.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[low] It only throws while the tracked run is returned (READY/RUNNING/SUCCEEDED). Otherwise it's silently replaced, which the "allowed when the tracked run is replaced" test pins down. Suggest "...throws while the tracked run is returned."

Comment thread src/actor.ts
static #instance?: Actor;

/**
* Tracks child runs of this Actor instance.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[nit] Restates the field name, can be dropped.

Comment thread src/utils.ts
}

/**
* A FIFO mutex that a critical section may re-enter from a nested call — the charge lock is taken by

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[nit] Doc still describes the charge lock / pushData(), but it's a shared util now (also per-name child-run locks), so make it general. Also the node:async_hooks import on L18 should go with the other builtins.

Comment thread src/child_run_tracker.ts
);
}

private async persist(trackedRuns: Record<string, TrackedChildRun>) {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[nit] persist() always gets this.loadedRuns, so the param is redundant. trackedRuns (promise) and loadedRuns also hold the same state, could be one field.

import log from '@apify/log';

import type { TrackedChildRun } from '../../src/child_run_tracker.js';
import { createIsolatedActor } from '../createIsolatedActor';

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[nit] Missing .js extension.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Holy guacamole, can we set up a lint rule for this?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I've set up a hook for myself months ago, so it didn't bother me, but sure, we can: #770

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

t-tooling Issues with this label are in the ownership of the tooling team. tested Temporary label used only programatically for some analytics.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Expose tracked child runs as a readable API Named child runs that survive migration and resurrection

5 participants