Skip to content

A single file of 2 GiB or more stops the scan, because project fingerprinting reads each file whole #266

Description

@webdevred

atom exits 1 before parsing anything when the scanned tree holds one file of 2 GiB or more. The AST cache fingerprint reads every regular file in the tree into a byte array to hash it, and Files.readAllBytes throws OutOfMemoryError: Required array size too large once a file will not fit in one array. No slice is written, and the .atom left behind is 8192 bytes where a run that gets through writes 16384 for the same two source files. The file does not have to be source: anything in the tree counts, a packfile, an archive, a database fixture.

Measured with atom 3.1.1, the Windows binary that cdxgen 13.2.0 bundles, on the Java 25.0.2 it ships with.

Reproduction

A directory with two files of JavaScript:

app/package.json

{
  "name": "atom-bigfile-repro",
  "version": "1.0.0",
  "private": true,
  "dependencies": { "lodash": "4.17.21" }
}

app/index.js

const lodash = require("lodash");

function greet(name) {
  return lodash.capitalize(name);
}

module.exports = { greet };

Put one more file beside them and run atom over the directory:

truncate -s 2147483648 app/big.pack
atom reachables -o app.atom -s reachables.json -l js app

The file is only ever hashed, never parsed, so truncate is enough and takes under a second. Remove app/.chen between runs, or the second run answers from the cache.

size of the extra file exit reachables slice
no extra file 0 written
2,147,483,647 (Integer.MAX_VALUE) 0 written
2,147,483,648 (2 GiB) 1 none

The threshold is per file and not per tree: the tree in the passing row is already over 2 GiB in total. It is an array length rather than an amount of free heap: the two runs either side of the boundary went back to back on the same machine, and the boundary sits exactly at Integer.MAX_VALUE.

Exception in thread "main" java.lang.OutOfMemoryError: Required array size too large
	at java.base@25.0.2/java.nio.file.Files.readAllBytes(Files.java:2979)
	at io.appthreat.x2cpg.passes.frontend.CpgCacheStore$.hashFile$$anonfun$1(CpgCacheStore.scala:65)
	at io.appthreat.x2cpg.passes.frontend.CpgCacheStore$.hashFile$$anonfun$adapted$1(CpgCacheStore.scala:66)
	at scala.util.Try$.apply(Try.scala:199)
	at io.appthreat.x2cpg.passes.frontend.CpgCacheStore$.hashFile(CpgCacheStore.scala:66)
	at io.appthreat.x2cpg.passes.frontend.CpgCacheStore$.projectFingerprint$$anonfun$1(CpgCacheStore.scala:53)
	[scala.runtime.function.JProcedure1 and scala.collection.immutable.List.foreach elided]
	at io.appthreat.x2cpg.passes.frontend.CpgCacheStore$.projectFingerprint(CpgCacheStore.scala:53)
	at io.appthreat.jssrc2cpg.utils.AstGenRunner.astGenFingerprint(AstGenRunner.scala:115)
	at io.appthreat.jssrc2cpg.utils.AstGenRunner.execute(AstGenRunner.scala:140)
	at io.appthreat.jssrc2cpg.JsSrc2Cpg.processAstGenDir$1(JsSrc2Cpg.scala:39)

The guard that is already there does not hold

hashFile wraps the read in a Try and logs a warning on failure, so a file it cannot read is meant to be tolerated. Try catches NonFatal only, and NonFatal.apply returns false for a VirtualMachineError, with a comment in the standard library saying so: "VirtualMachineError includes OutOfMemoryError and other fatal errors". The stack shows the throw passing through scala.util.Try$.apply and out of the process.

No option gets past it

Each of these still exits 1 with no slice on the same tree:

--no-ast-cache
--cache none
--exclude big.pack
--exclude-regex .*[.]pack

Two reasons in the code. astGenFingerprint calls projectFingerprint(config.inputPath, Seq(...)) with no third argument, so excludeDirNames stays at its default of the cache directory name alone and the caller's excludes never reach the walk. And in execute, val fp = astGenFingerprint is evaluated before cache.restore(fp, out.path), so the cache mode is consulted after the fingerprint has already been taken. The one path that skips it is the branch where config.astGenOutDir is non-empty and already holds astgen .json output, which a first run does not have.

What it costs when it does not crash

projectFingerprint keeps every regular file the walk finds, so the digest reads whatever else is in the tree, .git and node_modules included, none of which can change a JavaScript AST. With a 2,100,000,000 byte file beside them, the two JavaScript files above take 13 seconds. The same two files with that neighbour removed take 1.3 seconds.

Through cdxgen

cdxgen passes the directory it was given, which for a repository scan is the repository root, so a repository whose packfile has grown past 2 GiB hits this on every scan. What the operator sees is:

WARN: atom exited with status 1; the analysis may be incomplete.
atom reported a failure and produced no reachables slice for js.
Unable to generate reachables slice using atom. Check if this is a supported language.

The last line points at the language, which is not what went wrong. The usages slice attempted next then fails in OdbStorage.verifyStorageVersion, on the .atom the crashed run left behind.

Suggestion

Hashing through a stream instead of Files.readAllBytes removes the ceiling and the peak memory together. Two smaller things beside it: skipping .git and node_modules in projectFingerprint would keep them out of the digest, and catching Throwable rather than relying on Try would let the existing guard hold for the failure it is there for.


I am happy to create a PR for this.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions