Skip to content

fix(flows): avoid StackOverflowError after many steps in BaseLlmFlow - #1565

Open
innoprej wants to merge 1 commit into
google:mainfrom
innoprej:fix/llm-flow-step-loop
Open

innoprej wants to merge 1 commit into
google:mainfrom
innoprej:fix/llm-flow-step-loop

Conversation

@innoprej

@innoprej innoprej commented Sep 27, 2026 •

Copy link
Copy Markdown

Please ensure you have read the contribution guide before creating a pull request.

Link to Issue or Description of Change

1. Link to an existing issue (if applicable):

2. Or, if no issue exists, describe the change:

Problem:

BaseLlmFlow.run subscribes to each step from inside the previous step's completion (concatWith → toList → flatMapPublisher → PersistBarrier.awaitPersisted(...).andThen(run(...))). When a step completes on the subscribing thread, the next step starts on the same call stack, 23 frames deeper. The built-in models and session services complete on the subscribing thread (Gemini through Flowable.fromFuture, Claude and the non-streaming LangChain4j and SpringAI paths through Flowable.just after a blocking call, the session services through Single.just or Single.fromCallable), so with a 1 MB thread stack, the default on x86-64, an agent that keeps calling tools fails with StackOverflowError after roughly 250–300 LLM calls, before the default maxLlmCalls of 500 can end the run. On aarch64 the default thread stack is 2 MB; on a 2 MB stack the same agent made 635–663 calls in my runs, so there the default limit usually comes first. The error never reaches onError: it is thrown out of subscribe(), or, when the run is subscribed on a scheduler thread, it goes to RxJavaPlugins as an UndeliverableException and the subscriber never receives a terminal signal (in my runs, blockingGet() returned only when a 15-second timeout fired). The issue has the measurements.

Solution:

Drive the steps with repeatUntil, as LoopAgent loops over its sub-agents on its resumable path:

  • run(InvocationContext) now runs one step per subscription of an inner Flowable.defer(...) and lets repeatUntil subscribe again until the flow ends. When a step completes synchronously, repeatUntil resubscribes from a loop instead of from the completion callback (FlowableRepeatUntil.RepeatSubscriber.subscribeNext in RxJava 3.1.12), so every step starts at the same stack depth. This is the Java counterpart of the while True: loop in adk-python's BaseLlmFlow.run_async.
  • The private method that runs one step is renamed from run to runStep, since it no longer runs the rest of the flow, and gets a short Javadoc. Its end-of-flow checks, the pause on a pending long-running call and the maxSteps cut-off are unchanged. Where it used to subscribe to the next step, it now sets a continueFlow flag and ends the step with the PersistBarrier wait.
  • The step counter and the flag are created per subscription, inside the outer Flowable.defer.

Events are still emitted as each step produces them, the next step still starts only after the Runner has persisted the previous step's events, and an error still ends the run. One difference: subscribing twice to the same Flowable returned by run(ctx) now runs the flow twice from the start, where the old code replayed the first step from its cache() and ran the later steps again. Agents do not do this: BaseAgent calls runAsyncImpl again for every subscription (through Flowable.defer), and LlmAgent.runAsyncImpl calls run each time.

Testing Plan

Unit Tests:

  • I have added or updated unit tests for my change.
  • All unit tests pass locally. (core; one Windows-only failure that is also on main, see the table below)

BaseLlmFlowTest.run_modelRespondingOnSubscribingThread_reachesMaxLlmCallsWithoutStackOverflow runs the flow with a model that returns a function call through Flowable.just every time and maxLlmCalls set to 2,000, and expects the run to end with LlmCallsLimitExceededException after 2,000 requests. It runs the flow without a Runner, so the history does not grow, and it takes about a second. Like PersistBarrierTest.largeStep_awaitsAllWithoutStackOverflow from #1336, it depends on the thread stack size: on main it fails with any stack from 256 KB up to 4 MB, which includes the 2 MB aarch64 default, while an 8 MB stack would let the old code reach 2,000 steps.

Windows 11 on x86-64, Microsoft OpenJDK 17.0.19, Maven 4.0.0-rc-3 via mvnw:

Run main (BaseLlmFlow.java from main, with the new test) This change
./mvnw -pl core test -Dtest=BaseLlmFlowTest 28 run, 1 error: the new test, java.lang.StackOverflowError 28 run, 0 failures
./mvnw -pl core test -Dmaven.test.failure.ignore=true not run default-test and basic each run 1884 tests, 24 skipped, 1 failure: LocalSkillSourceTest.testListResources, a Windows-only failure that is also on main (#1541; fix in #1542). Without the flag the build stops at that failure in default-test. Of the other four surefire executions, three pass and apigee-llm-proxy-url runs no tests (#1557).

Manual End-to-End (E2E) Tests:

Not tested against a live model. The reproduction in the issue runs an LlmAgent with an EchoTool in an InMemoryRunner, with a TestLlm that asks for the tool on every call and answers on the subscribing thread:

main This change
1 MB thread stack (the x86-64 default) StackOverflowError after 253–265 LLM calls (14 runs) LlmCallsLimitExceededException after 500
-Xss256k StackOverflowError after 42 LlmCallsLimitExceededException after 500
-Xss2m (the aarch64 default size) stops at the 500-call limit; with maxLlmCalls raised to 2,000, StackOverflowError after 635–663 calls LlmCallsLimitExceededException at the limit (500, or 2,000 when raised)
Run subscribed with .subscribeOn(Schedulers.io()) and .timeout(15, SECONDS) stack overflow after 254–263 calls, an UndeliverableException to RxJavaPlugins, no signal until the timeout fires LlmCallsLimitExceededException after 500, in under 2 s

Recording the StackWalker depth at each model response shows 23 more frames per step on main, and the same depth at every step with this change, also over 20,000 steps on a 256 KB stack.

Checklist

  • I have read the CONTRIBUTING.md document.
  • My pull request contains a single commit.
  • I have performed a self-review of my own code.
  • I have commented my code, particularly in hard-to-understand areas.
  • I have added tests that prove my fix is effective or that my feature works.
  • New and existing unit tests pass locally with my changes. (core; see the table for the one failure that is also on main)
  • I have manually tested my changes end-to-end.
  • Any dependent changes have been merged and published in downstream modules.

Additional context

Rebased onto main at 2d4a59e, which already imports AtomicBoolean and adds the legacyResumption check that runStep keeps. The only conflict was the AtomicInteger import.

@hemasekhar-p hemasekhar-p self-assigned this Sep 28, 2026
@hemasekhar-p

Copy link
Copy Markdown
Contributor

Hi @innoprej, thank you for your contribution We appreciate you taking the time to submit this pull request. Currently this PR is under review by our team we will keep you posted if any additional information is required. thank you.

Comment thread core/src/main/java/com/google/adk/flows/llmflows/BaseLlmFlow.java Outdated
Comment thread core/src/main/java/com/google/adk/flows/llmflows/BaseLlmFlow.java Outdated
@MiloszSobczyk

Copy link
Copy Markdown
Member

Helllo, could you please merge the newest changes that were introduced in the code? Afterwards, please ensure that the code still works correctly.

BaseLlmFlow.run subscribed to the next step from inside the previous
step's completion, through concatWith, toList, flatMapPublisher and
PersistBarrier.awaitPersisted(...).andThen(run(...)). When a step
completed on the subscribing thread, as it does with the core models
and the session services, each step ran 23 stack frames deeper than
the one before. With a 1 MB thread stack, the default on x86-64, an
agent that kept calling tools failed with StackOverflowError after
roughly 250-300 LLM calls, before the default maxLlmCalls of 500
could end the run, and the error never reached onError.

Drive the steps with repeatUntil instead, as LoopAgent does on its
resumable path. repeatUntil resubscribes from a loop when the source
completes synchronously, so every step starts at the same stack
depth, like the while loop in adk-python's BaseLlmFlow.run_async.
The end-of-flow checks, the pause on a pending long-running call,
the PersistBarrier wait between steps and maxSteps are unchanged.
@innoprej
innoprej force-pushed the fix/llm-flow-step-loop branch from 5d92339 to 7470b43 Compare September 30, 2026 21:15
@innoprej

Copy link
Copy Markdown
Author

Done. I rebased the single commit onto the current main (2d4a59e). The only conflict was in the imports: main now imports AtomicBoolean itself, so this change only adds AtomicInteger; the rest of the diff is unchanged. main's new resumption logic is kept as is: StepResume runs inside runOneStep, which this change doesn't touch, and the legacyResumption pause in runStep returns without setting continueFlow, so the loop ends there, as it does at the other end conditions.

On Windows 11 with JDK 17, BaseLlmFlowTest, PersistBarrierTest, RunnerResumabilityTest and RunnerLegacyResumabilityTest pass (95 tests in both the default-test and basic executions). With only the test file applied to the same main, the new regression test still fails with StackOverflowError. The full core suite (2,016 tests in each of those executions) has one failure, LocalSkillSourceTest.testListResources, the Windows-only failure that is also on main and that #1542 fixes.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

BaseLlmFlow nests each step inside the previous one and overflows the stack after a few hundred LLM calls

3 participants