| #728 |
Missing rollback_pending() on savepoint release failure in sleep path
In _handle_sleep_exception(), if RELEASE SAVEPOINT fails, pending cache entries are never rolled bac...
|
closed |
medium |
2026-02-07 02:48 |
- |
|
| #727 |
No circular dependency detection in workflow DAG
Circular deps in workflow graph cause infinite hang with zero error output. Need cycle detection usi...
|
closed |
medium |
2026-02-07 02:48 |
- |
|
| #726 |
On-failure callback double execution risk
In _handle_failure_exception(), task is marked as executed AFTER callback runs. If callback raises, ...
|
closed |
high |
2026-02-07 02:48 |
- |
|
| #725 |
State machine guard on workflow_run status updates
update_workflow_run_status() updates status without checking current state. Two concurrent writers c...
|
closed |
high |
2026-02-07 02:48 |
- |
|
| #724 |
## Problem
Two issues with worker signal handling prevent proper graceful shutdown:
### Issue 1: Delayed Shutdown Response
- Signal handler sets `_shutdown_requested = True`
- But `wakeup_event` is a local variable in `worker_command()`
- Signal handler can't access it to wake up blocked `Event.wait()`
- Worker waits up to `poll_interval` (1-5s) before noticing SIGTERM
### Issue 2: In-Flight Tasks Abandoned
- After main loop breaks, `finally` block stops services immediately
- No wait for tasks in `orchestrator.bulkhead` to complete
- Tasks get killed mid-execution
- They stay claimed until heartbeat timeout, then re-execute from scratch
## Impact
- K8s rolling updates will kill tasks mid-execution
- Wasted compute from task re-execution
- Potential data corruption for non-idempotent tasks
## Solution
1. Make `wakeup_event` module-level so signal handler can access it
2. Add graceful drain logic like activity_worker has:
- Wait up to 30s for in-flight tasks
- Call `bulkhead.shutdown(wait=True)`
## Files
- engine/cli/worker.py
## Testing
- Send SIGTERM to worker with active tasks
- Verify tasks complete before worker exits
- Verify shutdown happens within seconds, not waiting for poll_interval
|
closed |
high |
2025-12-28 00:21 |
- |
|
| #722 |
Refactor wait_for_parallel_branches to use JoinMode for consistency
Currently wait_for_parallel_branches uses ad-hoc fail_on_error=True boolean while JoinOperator has p...
|
closed |
high |
2025-12-27 21:09 |
- |
|
| #721 |
Workers become zombies after asyncio crash - container healthy but all threads dead
OBSERVED: Workers can crash with asyncio.exceptions.CancelledError during HTTP operations, causing a...
|
closed |
critical |
2025-12-27 03:27 |
- |
|
| #720 |
Sandbox temp files fail in Docker-in-Docker: can't find __main__ module
FIXED: When running sandboxed code execution (tools.code.exec) inside Docker containers with Docker ...
|
closed |
high |
2025-12-27 03:27 |
- |
|
| #719 |
Parallel branch retry creates orphaned waits due to absurd_run_id in join_event_name
FIXED: When a workflow with parallel branches is retried (e.g., due to worker crash), the join_event...
|
closed |
critical |
2025-12-27 02:52 |
- |
|
| #718 |
Branch tasks now inherit retry policy from ParallelOperator
Fixed branch execution tasks to inherit max_attempts and retry_strategy from the ParallelOperator's ...
|
closed |
high |
2025-12-26 23:58 |
- |
|
| #717 |
Branch tasks now retry on worker crash
Fixed branch execution tasks to have max_attempts=3 with exponential backoff (10s, 30s). Previously ...
|
closed |
high |
2025-12-26 23:53 |
- |
|
| #716 |
Skip propagation bug in switch/condition operators
When switch/condition marks unselected branches as executed, downstream tasks depending on skipped b...
|
closed |
high |
2025-12-26 16:17 |
- |
|
| #715 |
BUG: python_task.py swallows AbsurdSleepError in sandbox fallback path
In engine/tools/python_task.py lines 373-376, the except Exception block catches AbsurdSleepError wh...
|
closed |
critical |
2025-12-26 00:06 |
- |
|
| #714 |
Add generic configurable update handler tool
Current state: The tools.testing.update_handler is hardcoded to handle 'approve_request' updates wit...
|
closed |
high |
2025-12-25 21:46 |
- |
|
| #713 |
ReflexiveOperator completes workflow before activity finishes
ReflexiveOperator calls execute_activity_operator which is queue-only, but doesn't wait for the acti...
|
closed |
high |
2025-12-25 21:02 |
- |
|
| #711 |
CFG-01: No schema validation for configuration
Location: engine/config.py. Issue: ConfigParser used without schema validation. No type/range/requir...
|
closed |
high |
2025-12-25 02:56 |
- |
|
| #710 |
VAL-05: JoinOperator join_tasks not validated
Location: highway_dsl/workflow_dsl.py. Issue: JoinOperator.join_tasks references not validated again...
|
closed |
high |
2025-12-25 02:56 |
- |
|
| #708 |
RETRY-01: No jitter in Absurd retry delays
Location: a_absurd.sql:639-654. Issue: Retry delays calculated without jitter. Causes thundering her...
|
closed |
high |
2025-12-25 02:56 |
- |
|
| #707 |
ERR-02: Non-deterministic time.sleep() inside transaction
Location: operators.py:247-262. Issue: Inline retry uses time.sleep() holding DB connection open. Vi...
|
closed |
high |
2025-12-25 02:56 |
- |
|
| #706 |
ERR-01: _fail_run exceptions logged but not propagated
Location: orchestrator.py:1198-1199. Issue: Fatal failure to log a run failure only logged, not esca...
|
closed |
high |
2025-12-25 02:56 |
- |
|