atom exits 1 before parsing anything when the scanned tree holds one file of 2 GiB or more. The AST cache fingerprint reads every regular file in the tree into a byte array to hash it, and Files.readAllBytes throws OutOfMemoryError: Required array size too large once a file will not fit in one array. No slice is written, and the .atom left behind is 8192 bytes where a run that gets through writes 16384 for the same two source files. The file does not have to be source: anything in the tree counts, a packfile, an archive, a database fixture.
Measured with atom 3.1.1, the Windows binary that cdxgen 13.2.0 bundles, on the Java 25.0.2 it ships with.
Reproduction
A directory with two files of JavaScript:
app/package.json
{
"name": "atom-bigfile-repro",
"version": "1.0.0",
"private": true,
"dependencies": { "lodash": "4.17.21" }
}
app/index.js
const lodash = require("lodash");
function greet(name) {
return lodash.capitalize(name);
}
module.exports = { greet };
Put one more file beside them and run atom over the directory:
truncate -s 2147483648 app/big.pack
atom reachables -o app.atom -s reachables.json -l js app
The file is only ever hashed, never parsed, so truncate is enough and takes under a second. Remove app/.chen between runs, or the second run answers from the cache.
| size of the extra file |
exit |
reachables slice |
| no extra file |
0 |
written |
2,147,483,647 (Integer.MAX_VALUE) |
0 |
written |
| 2,147,483,648 (2 GiB) |
1 |
none |
The threshold is per file and not per tree: the tree in the passing row is already over 2 GiB in total. It is an array length rather than an amount of free heap: the two runs either side of the boundary went back to back on the same machine, and the boundary sits exactly at Integer.MAX_VALUE.
Exception in thread "main" java.lang.OutOfMemoryError: Required array size too large
at java.base@25.0.2/java.nio.file.Files.readAllBytes(Files.java:2979)
at io.appthreat.x2cpg.passes.frontend.CpgCacheStore$.hashFile$$anonfun$1(CpgCacheStore.scala:65)
at io.appthreat.x2cpg.passes.frontend.CpgCacheStore$.hashFile$$anonfun$adapted$1(CpgCacheStore.scala:66)
at scala.util.Try$.apply(Try.scala:199)
at io.appthreat.x2cpg.passes.frontend.CpgCacheStore$.hashFile(CpgCacheStore.scala:66)
at io.appthreat.x2cpg.passes.frontend.CpgCacheStore$.projectFingerprint$$anonfun$1(CpgCacheStore.scala:53)
[scala.runtime.function.JProcedure1 and scala.collection.immutable.List.foreach elided]
at io.appthreat.x2cpg.passes.frontend.CpgCacheStore$.projectFingerprint(CpgCacheStore.scala:53)
at io.appthreat.jssrc2cpg.utils.AstGenRunner.astGenFingerprint(AstGenRunner.scala:115)
at io.appthreat.jssrc2cpg.utils.AstGenRunner.execute(AstGenRunner.scala:140)
at io.appthreat.jssrc2cpg.JsSrc2Cpg.processAstGenDir$1(JsSrc2Cpg.scala:39)
The guard that is already there does not hold
hashFile wraps the read in a Try and logs a warning on failure, so a file it cannot read is meant to be tolerated. Try catches NonFatal only, and NonFatal.apply returns false for a VirtualMachineError, with a comment in the standard library saying so: "VirtualMachineError includes OutOfMemoryError and other fatal errors". The stack shows the throw passing through scala.util.Try$.apply and out of the process.
No option gets past it
Each of these still exits 1 with no slice on the same tree:
--no-ast-cache
--cache none
--exclude big.pack
--exclude-regex .*[.]pack
Two reasons in the code. astGenFingerprint calls projectFingerprint(config.inputPath, Seq(...)) with no third argument, so excludeDirNames stays at its default of the cache directory name alone and the caller's excludes never reach the walk. And in execute, val fp = astGenFingerprint is evaluated before cache.restore(fp, out.path), so the cache mode is consulted after the fingerprint has already been taken. The one path that skips it is the branch where config.astGenOutDir is non-empty and already holds astgen .json output, which a first run does not have.
What it costs when it does not crash
projectFingerprint keeps every regular file the walk finds, so the digest reads whatever else is in the tree, .git and node_modules included, none of which can change a JavaScript AST. With a 2,100,000,000 byte file beside them, the two JavaScript files above take 13 seconds. The same two files with that neighbour removed take 1.3 seconds.
Through cdxgen
cdxgen passes the directory it was given, which for a repository scan is the repository root, so a repository whose packfile has grown past 2 GiB hits this on every scan. What the operator sees is:
WARN: atom exited with status 1; the analysis may be incomplete.
atom reported a failure and produced no reachables slice for js.
Unable to generate reachables slice using atom. Check if this is a supported language.
The last line points at the language, which is not what went wrong. The usages slice attempted next then fails in OdbStorage.verifyStorageVersion, on the .atom the crashed run left behind.
Suggestion
Hashing through a stream instead of Files.readAllBytes removes the ceiling and the peak memory together. Two smaller things beside it: skipping .git and node_modules in projectFingerprint would keep them out of the digest, and catching Throwable rather than relying on Try would let the existing guard hold for the failure it is there for.
I am happy to create a PR for this.
atomexits 1 before parsing anything when the scanned tree holds one file of 2 GiB or more. The AST cache fingerprint reads every regular file in the tree into a byte array to hash it, andFiles.readAllBytesthrowsOutOfMemoryError: Required array size too largeonce a file will not fit in one array. No slice is written, and the.atomleft behind is 8192 bytes where a run that gets through writes 16384 for the same two source files. The file does not have to be source: anything in the tree counts, a packfile, an archive, a database fixture.Measured with atom 3.1.1, the Windows binary that cdxgen 13.2.0 bundles, on the Java 25.0.2 it ships with.
Reproduction
A directory with two files of JavaScript:
app/package.json{ "name": "atom-bigfile-repro", "version": "1.0.0", "private": true, "dependencies": { "lodash": "4.17.21" } }app/index.jsPut one more file beside them and run atom over the directory:
The file is only ever hashed, never parsed, so
truncateis enough and takes under a second. Removeapp/.chenbetween runs, or the second run answers from the cache.Integer.MAX_VALUE)The threshold is per file and not per tree: the tree in the passing row is already over 2 GiB in total. It is an array length rather than an amount of free heap: the two runs either side of the boundary went back to back on the same machine, and the boundary sits exactly at
Integer.MAX_VALUE.The guard that is already there does not hold
hashFilewraps the read in aTryand logs a warning on failure, so a file it cannot read is meant to be tolerated.TrycatchesNonFatalonly, andNonFatal.applyreturns false for aVirtualMachineError, with a comment in the standard library saying so: "VirtualMachineError includes OutOfMemoryError and other fatal errors". The stack shows the throw passing throughscala.util.Try$.applyand out of the process.No option gets past it
Each of these still exits 1 with no slice on the same tree:
Two reasons in the code.
astGenFingerprintcallsprojectFingerprint(config.inputPath, Seq(...))with no third argument, soexcludeDirNamesstays at its default of the cache directory name alone and the caller's excludes never reach the walk. And inexecute,val fp = astGenFingerprintis evaluated beforecache.restore(fp, out.path), so the cache mode is consulted after the fingerprint has already been taken. The one path that skips it is the branch whereconfig.astGenOutDiris non-empty and already holds astgen.jsonoutput, which a first run does not have.What it costs when it does not crash
projectFingerprintkeeps every regular file the walk finds, so the digest reads whatever else is in the tree,.gitandnode_modulesincluded, none of which can change a JavaScript AST. With a 2,100,000,000 byte file beside them, the two JavaScript files above take 13 seconds. The same two files with that neighbour removed take 1.3 seconds.Through cdxgen
cdxgen passes the directory it was given, which for a repository scan is the repository root, so a repository whose packfile has grown past 2 GiB hits this on every scan. What the operator sees is:
The last line points at the language, which is not what went wrong. The usages slice attempted next then fails in
OdbStorage.verifyStorageVersion, on the.atomthe crashed run left behind.Suggestion
Hashing through a stream instead of
Files.readAllBytesremoves the ceiling and the peak memory together. Two smaller things beside it: skipping.gitandnode_modulesinprojectFingerprintwould keep them out of the digest, and catchingThrowablerather than relying onTrywould let the existing guard hold for the failure it is there for.I am happy to create a PR for this.