Skip to content

Reduce memory consumption of large literal arrays in 2.3.x - #6727

Open
alfredbez wants to merge 1 commit into
phpstan:2.3.xfrom
alfredbez:fix/literal-array-memory-2.3
Open

alfredbez wants to merge 1 commit into
phpstan:2.3.xfrom
alfredbez:fix/literal-array-memory-2.3

Conversation

@alfredbez

Copy link
Copy Markdown
Contributor

When analysing files with large literal arrays, PHPStan 2.3.x consumed ~55% more peak memory than 2.2.16 (139 MB vs 77 MB on 20,000 items).

Memory consumption on reproducer (php run.php [entries]):

Entries 2.2.16 2.3.x before 2.3.x with fix
5,000 16 MB 26 MB 16 MB
10,000 34 MB 62 MB 34 MB
20,000 77 MB (2.28 s) 139 MB (2.66 s) 76 MB (2.26 s)
40,000 168 MB ~285 MB 166 MB (6.21 s)

What caused the increase:

  1. In ScalarHandler, each scalar literal allocated an InitializerExprContext, an 800-byte closure (typeCallback), and an ExpressionResult. For 20,000 items (40,000 scalar expressions), closures and context objects alone consumed over 38 MB. Precomputing constant scalar types (ConstantStringType, ConstantIntegerType, ConstantFloatType) directly avoids allocating closures and contexts.
  2. ArrayHandler accumulated all key and value results into $itemResults and captured them in the array's typeCallback closure. Scalar item types are constant and can be created on demand in the callback, so $itemTypes only records non-scalar items.
  3. For arrays without by-reference items, ArrayHandler now evicts the item's key and value from ExpressionResultStorage once the item's callback has finished. This prevents 40,000 ExpressionResult instances from accumulating in SplObjectStorage throughout the loop.
  4. AssignHandler invoked processArrayByRefItems on every assigned array literal, even when the array contained no by-reference elements. Guarding with hasArrayReference() avoids traversing large arrays during assignment, and lookups fall back to $scope->getType() if a stored result is absent.
  5. DuplicateKeysInLiteralArraysRule eagerly pretty-printed all keys via exprPrinter and stored them in $printedValues. Storing only the first node occurrence and formatting printed representations lazily when a duplicate key is actually detected avoids allocating tens of thousands of printed strings for unique keys.

All 22,299 tests pass.

@ondrejmirtes

Copy link
Copy Markdown
Member

Hey, please submit an issue with the description and reproducer first, thank you.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants