Repository navigation
feat(durable-jobs): add grain-scoped feature handlers - #10679
Conversation
There was a problem hiding this comment.
Pull request overview
This PR evolves the DurableJobs execution model to support activation-scoped (framework-owned) job handlers alongside existing grain-owned handlers, while also distinguishing successful durable reschedules from failure-driven retries. It updates the runtime execution pipeline, shard persistence APIs, and receiver-side deduplication semantics, with corresponding public API surface updates and new/updated tests.
Changes:
- Added activation-scoped feature handler registration/dispatch via
IDurableJobHandlerRegistry+IDurableJobFeatureHandler, with ordinal job-name matching and precedence overIDurableJobHandler. - Introduced explicit
RetryAt/durable-reschedule disposition (DurableJobRunStatus.RetryAt+DurableJobRunResult.RetryAt(...)) and a newIJobShard.RescheduleJobAsync(...)API to persist a dequeue-count reset distinct from failure retry. - Made completed-execution deduplication identity
(JobId, RunId)and added a configurable completed-attempt retention horizon.
Show a summary per file
| File | Description |
|---|---|
| test/Orleans.DurableJobs.Tests/DurableJobs/JobShardTests.cs | New unit coverage for reschedule vs retry dequeue-count persistence behavior. |
| test/Orleans.DurableJobs.Tests/DurableJobs/JobShardManagerTestsRunner.cs | Adds reassignment scenario coverage for reschedule dequeue-count reset persistence. |
| test/Orleans.DurableJobs.Tests/DurableJobs/DurableJobFeatureHandlerTests.cs | New tests for handler registry ordinal behavior and serialization/status compatibility. |
| test/Orleans.Core.Tests/DurableJobs/ShardExecutorTests.cs | Adds executor tests for RetryAt and unknown disposition handling. |
| test/Orleans.Core.Tests/DurableJobs/LocalDurableJobManagerTests.cs | Updates test shards to implement the new RescheduleJobAsync API. |
| test/Orleans.Core.Tests/DurableJobs/DurableJobsExtensionsTests.cs | New test verifying DI scoping/sharing of registry and receiver extension. |
| test/Orleans.Core.Tests/DurableJobs/DurableJobReceiverExtensionTests.cs | Updates dedup identity to (JobId, RunId) and adds feature-handler precedence + retention tests. |
| src/Orleans.DurableJobs/ShardExecutor.cs | Executes RetryAt by calling shard reschedule and treats unknown dispositions via failure-retry policy. |
| src/Orleans.DurableJobs/JournaledJobShard.cs | Adds durable RescheduleJobAsync and persists reset dequeue count through journaled state. |
| src/Orleans.DurableJobs/JobShard.cs | Adds IJobShard.RescheduleJobAsync and implements dequeue-count reset semantics. |
| src/Orleans.DurableJobs/IDurableJobReceiverExtension.cs | Adds activation-scoped feature handler dispatch and updates dedup to (JobId, RunId) with retention option. |
| src/Orleans.DurableJobs/IDurableJobHandlerRegistry.cs | Introduces public registry + feature handler APIs and internal registry implementation. |
| src/Orleans.DurableJobs/Hosting/DurableJobsOptions.cs | Adds CompletedJobAttemptRetentionPeriod option and validation. |
| src/Orleans.DurableJobs/Hosting/DurableJobsExtensions.cs | Wires the activation-scoped registry into DI and receiver extension construction. |
| src/Orleans.DurableJobs/DurableJobRunResult.cs | Adds RetryAtTime + IsRetryRequested and the RetryAt disposition/status value. |
| src/api/Orleans.DurableJobs/Orleans.DurableJobs.cs | Updates generated public API surface for new/changed DurableJobs APIs. |
Review details
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
- Files reviewed: 16/16 changed files
- Comments generated: 1
- Review effort level: Lite
There was a problem hiding this comment.
Review details
Suppressed comments (2)
Previously missed (1) — in code that hasn't changed since the last review.
src/Orleans.DurableJobs/ShardExecutor.cs:286
- The
RetryAtpath represents a successful handler outcome (durable reschedule), but the activity status is not set toOklike it is forCompleted. Setting it explicitly would make traces consistent with other success outcomes.
LogRetryingJob(_logger, jobContext.Job.Id, jobContext.Job.Name, retryAt, jobContext.DequeueCount);
await resettableShard.RescheduleJobAsync(jobContext, retryAt, cancellationToken);
_durableJobsInstruments.OnJobRetried(_timeProvider.GetElapsedTime(attemptStartTimestamp));
activity?.SetTag(ActivityTagKeys.DurableJobStatus, "rescheduled");
break;
src/Orleans.DurableJobs/ShardExecutor.cs:285
- In the
RetryAt(successful durable reschedule) path, this still callsLogRetryingJoband_durableJobsInstruments.OnJobRetried(...), which will classify successful reschedules as "retried" in logs/metrics. That undermines the stated goal of distinguishing successful rescheduling from failure retries for accounting/observability. Consider introducing a separate log/metric path (e.g., "rescheduled") or at minimum avoiding the "retried" counter/tag forRetryAt.
LogRetryingJob(_logger, jobContext.Job.Id, jobContext.Job.Name, retryAt, jobContext.DequeueCount);
await resettableShard.RescheduleJobAsync(jobContext, retryAt, cancellationToken);
_durableJobsInstruments.OnJobRetried(_timeProvider.GetElapsedTime(attemptStartTimestamp));
activity?.SetTag(ActivityTagKeys.DurableJobStatus, "rescheduled");
- Files reviewed: 15/15 changed files
- Comments generated: 3
- Review effort level: Lite
There was a problem hiding this comment.
Review details
Suppressed comments (1)
Previously missed (1) — in code that hasn't changed since the last review.
src/Orleans.DurableJobs/DurableJobRunResult.cs:28
IsRetryRequestedis annotated withMemberNotNullWhen(true, nameof(RetryAtTime)), but it currently returns true purely based onStatus == RetryAt. This is not guaranteed to implyRetryAtTimeis non-null (e.g., a malformed/deserialized/reflectively-constructed result can haveStatus=RetryAtandRetryAtTime=null, which this PR explicitly handles elsewhere). Make the property consistent with its nullability contract by also checkingRetryAtTime.
/// <summary>
/// Gets a value indicating whether the job requested durable rescheduling.
/// </summary>
[MemberNotNullWhen(true, nameof(RetryAtTime))]
public bool IsRetryRequested => Status == DurableJobRunStatus.RetryAt;
- Files reviewed: 18/18 changed files
- Comments generated: 0 new
- Review effort level: Lite
There was a problem hiding this comment.
Review details
Suppressed comments (1)
Previously missed (1) — in code that hasn't changed since the last review.
src/Orleans.DurableJobs/JobShard.cs:203
JobRunContext's third constructor parameter is namedretryCountbut it setsDequeueCount. Using a named argument here makes the code harder to read (it looks like a different concept than the dequeue count you're intentionally resetting). Consider passing the value positionally to avoid coupling to a misleading parameter name.
var resetContext = new JobRunContext(jobContext.Job, jobContext.RunId, retryCount: 0);
- Files reviewed: 18/18 changed files
- Comments generated: 0 new
- Review effort level: Lite
There was a problem hiding this comment.
Review details
Suppressed comments (1)
Previously missed (1) — in code that hasn't changed since the last review.
src/Orleans.DurableJobs/IDurableJobReceiverExtension.cs:105
HandlerExecutionTrackerrequires exactly one terminal call (Completed/Canceled/Failed), butExecuteFeatureHandlerAsynccallstracker.Completed()before validating the returned result. If a feature handler returnsnull, the subsequent exception will be caught andtracker.Failed(exception)will also run, double-counting metrics and violating the tracker contract.
var result = await handler.ExecuteJobAsync(context, cancellationToken);
tracker.Completed();
return result ?? throw new InvalidOperationException(
$"Durable job feature handler for '{context.Job.Name}' returned a null result.");
- Files reviewed: 20/20 changed files
- Comments generated: 0 new
- Review effort level: Lite
c0e879c to
bf0051a
Compare
Problem
DurableJobs currently requires the grain itself to own every job handler. Framework features need activation-local dispatch without changing the grain's public contract. Cancellation also needs distinct semantics for the durable task and its current execution attempt.
Solution
IDurableJobHandlerRegistryandIDurableJobFeatureHandler. Each feature handler is registered once and owns deterministic, side-effect-free name matching throughCanHandle(string jobName), supporting exact names, prefixes, or regexes.IDurableJobHandleras the safe grain-author contract: successful return completes the job and exceptions follow retry policy. Feature handlers are the advanced contract which returns an explicitDurableJobRunResult. Both contracts flow through one centralized cancellation, exception, result, metrics, and tracing path.IDurableJobHandler; overlapping feature matches fail explicitly as ambiguous. Job names are validated before persistence at manager and shard boundaries.(JobId, ExecutionGeneration, DequeueCount), which remains stable across duplicate delivery even whenRunIdchanges. Concurrent deliveries share the active invocation; once it reaches a terminal disposition, a later delivery starts a new invocation.InProgressas non-terminal: duplicate calls share the current invocation, and the feature handler is invoked again after its requested delay.CancelAsyncas a durable task cancellation request. Atrueresult means removal was durably applied, preventing future attempts; an already-running attempt may still complete cooperatively.attemptCancellationTokenas cooperative cancellation of only the current attempt. Host shutdown or ownership loss stops that attempt while preserving the task for reassignment. Active long polls end promptly on the request while a handler which has not stopped remains active.DurableJobMutationResultvalues prevent suppressed or unowned mutations from being reported as successful.attempt_canceledin shard metrics, and is identified explicitly in logs and mutation activities.Rationale
The simple grain contract makes completion and failure difficult to misuse. The advanced feature contract exposes polling and rescheduling only to components which deliberately opt into the DurableJobs disposition model. Feature handlers compose with existing execution and retry behavior while claiming grain-scoped work using stable name-matching rules. Durable task cancellation governs future delivery; attempt cancellation governs only the current host execution. Handlers own any application-specific idempotency needed when a terminal invocation is delivered again.