Skip to content

usages: whole types silently dropped from userDefinedTypes when several files are scanned together (c) #258

Description

@mquiroga42

Summary

atom usages -l c silently drops whole entries from userDefinedTypes in the emitted
usages.json. atom exits 0, prints no warning, writes well-formed JSON, and lists every scanned
file in objectSlices — the types are simply not there.

Which types go missing changes from run to run over a byte-identical tree.

Reproduced on v3.1.1 (current release) and v3.0.3, Linux 5.15 (WSL2), x86-64, 8 cores, with
the bundled JDK 23.

Reproduction

Two files per module: a uniquely-named header with one named struct typedef, and a .c that includes
it and uses the type. No build system, no external headers.

#!/bin/bash
# make-tree.sh <root> [n]
set -eu
root="$1"; n="${2:-10}"; rm -rf "$root"
for i in $(seq 1 "$n"); do
  d="$root/mod$i"; mkdir -p "$d"
  cat > "$d/Cfg$i.h" <<EOF
#ifndef CFG_${i}_H
#define CFG_${i}_H
typedef struct Tipo$i { int f${i}a; long f${i}b; } Tipo$i;
#endif
EOF
  cat > "$d/usa$i.c" <<EOF
#include "Cfg$i.h"
int usa$i(void){ Tipo$i v; v.f${i}a = $i; v.f${i}b = $i; return (int)(v.f${i}a + v.f${i}b); }
EOF
done
./make-tree.sh /tmp/tree 10
atom usages -o /tmp/app.atom -s /tmp/usages.json -l c /tmp/tree
python3 -c "
import json
names = {u['name'] for u in json.load(open('/tmp/usages.json'))['userDefinedTypes']}
print(sorted(f'Tipo{i}' for i in range(1,11) if f'Tipo{i}' not in names))"

Rebuild the tree for every run. atom writes its .chen/ cache inside the directory it scans, so
re-running over the same tree replays the previous answer and looks deterministic when it is not.

Observed

Expected: all ten types present. Actual, on v3.1.1, 14 consecutive runs, fresh tree each time:

run  1: missing Tipo1, Tipo3, Tipo5, Tipo6
run  2: missing Tipo1, Tipo3, Tipo6
run  3: missing Tipo3, Tipo6
run  4: missing Tipo1, Tipo3, Tipo6, Tipo10
run  5: missing Tipo1, Tipo4
run  6: missing Tipo1, Tipo3, Tipo6, Tipo7
run  7: missing Tipo3, Tipo5
run  8: missing Tipo3, Tipo10
run  9: missing Tipo4, Tipo6
run 10: missing Tipo1, Tipo3
run 11: missing Tipo1, Tipo4, Tipo6, Tipo7
run 12: missing Tipo1, Tipo4, Tipo6
run 13: missing Tipo1, Tipo6
run 14: missing Tipo1, Tipo3, Tipo6, Tipo7

14 of 14 runs lost at least one type, 2–4 per run.

Rate by tree size (v3.0.3, types missing per run):

modules runs result
2 5 0 0 0 0 0 — never
3 5 0 0 0 0 1 — starts here
5 14 13/14 runs lost ≥1 (1–2 each)
10 14 14/14 runs lost ≥1 (1–5 each)

One full run, verbatim:

$ atom usages -o app.atom -s verif.json -l c /tmp/tree
Auto-discovered 1 project include paths
Slicing the atom for usages. This might take a few minutes ...
Slices have been successfully written to /tmp/verif.json
$ echo $?
0

objectSlices covers all ten files:

mod1/usa1.c mod10/usa10.c mod2/usa2.c mod3/usa3.c mod4/usa4.c
mod5/usa5.c mod6/usa6.c mod7/usa7.c mod8/usa8.c mod9/usa9.c

userDefinedTypes holds six distinct types, each duplicated once:

Tipo8 Tipo8 Tipo9 Tipo9 Tipo7 Tipo7 Tipo6 Tipo6 Tipo2 Tipo2 Tipo10 Tipo10

Tipo1, Tipo3, Tipo4 and Tipo5 appear zero times anywhere in the raw JSON.

Narrowing it down

Each type is fine on its own. Scanning each module alone, one .c per run: 10 of 10 types
present, every time. Nothing here is intrinsically unparseable — the loss needs several files in one
run.

Header names are unique on purpose. With a colliding basename (every module shipping its own
Config.h) types also go missing, but that is the expected consequence of an ambiguous #include,
it is deterministic, and it is a different thing. Unique names remove that effect entirely.

Not uniform across types. Over 14 runs at 10 modules (v3.0.3):

Tipo1: 11/14   Tipo2: 0/14   Tipo3: 8/14   Tipo4: 6/14   Tipo5:  2/14
Tipo6: 10/14   Tipo7: 2/14   Tipo8: 0/14   Tipo9: 0/14   Tipo10: 5/14

Some types are lost most of the time, others never, with no obvious positional pattern. v3.1.1 shows
the same skew — Tipo1, Tipo3 and Tipo6 dominate there too.

Ruled out by measurement:

hypothesis test result
CPU parallelism taskset -c 0, 8 runs still loses 1–4
Flux engine fragment cache --legacy-dataflow, 8 runs still loses 3–5
memory pressure JAVA_OPTS=-Xmx8g, 6 runs still loses 2–4
stale .chen/ cache fresh tree every run still loses

Pinning to a single core still loses types, so it is not contention between worker threads on
separate cores.

Why it matters

usages.json is what cdxgen consumes for a source-code SBOM (--usages-slices-file). Types missing
from it are missing downstream, and because atom exits 0 with no diagnostic there is nothing to alert
on: the output looks complete.

Environment

atom      3.1.1 (also reproduced on 3.0.3)
JDK       23 (bundled)
OS        Linux 5.15.146.1-microsoft-standard-WSL2, x86-64, 8 cores
language  c

Happy to run further experiments against this repro if it helps narrow it down.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions