Describe the bug
pjlib-test intermittently aborts on a pool-alignment assertion when run with --shuffle and more than one worker thread. Seen on the macOS CI runners:
Shuffling tests, random seed=312
Performing 2 features tests with 2 worker threads
Assertion failed: (((alignment)>0) && ((alignment) & ((alignment)-1))==0),
function pj_pool_alloc_from_block, file pool_i.h, line 48.
Abort trap: 6
It aborts roughly 35 ms into the run, so before much has happened, and it is not reproducible on demand — the same commit passes the same job on other runs.
What the assertion means
The value cannot have come from the caller. pj_pool_aligned_alloc() already screens it (pjlib/include/pj/pool_i.h:81):
PJ_ASSERT_RETURN(!alignment || PJ_IS_POWER_OF_TWO(alignment), NULL);
if (!alignment)
alignment = pool->alignment;
A caller passing 0 — the ordinary "use the pool default" case — is accepted, and the value is then taken from pool->alignment and passed to pj_pool_alloc_from_block(), which asserts it is a power of two. So reaching that assertion means pool->alignment itself was 0, or otherwise not a power of two, at the moment of the call.
Two things follow, and they are separable:
1. The substituted value is never validated. The caller's alignment is screened on the line above; the pool's own is not. Whatever the underlying cause, pool->alignment becoming invalid currently surfaces as an assertion three frames away from wherever the pool was actually damaged. Validating it at pool creation — so the field cannot hold a non-power-of-two — or screening it at the substitution would turn this into a diagnosable failure instead of a puzzle.
2. Something is leaving pool->alignment invalid. I have not diagnosed this and do not want to guess. Zero is what an uninitialised or zeroed pool header would hold, and the failure only appears under --shuffle with two workers, which is suggestive of a lifetime or concurrency fault — a pool used after release, or one being created concurrently with use — rather than a bad call site. That is a hypothesis, not a finding.
Steps to reproduce
No reliable reproduction. It appears sporadically in CI under the existing shuffled multi-worker configuration:
pjlib-test-$(target) --ci-mode -w 2 --shuffle --stdout-buf 1 ssl_sock_test ssl_sock_stress_test
The seed is printed on each run (random seed=312 above), so a failing ordering should be replayable if the harness accepts a fixed seed — that seems the cheapest first step toward a reproduction, and is why I have quoted it.
Why I am filing it separately
It surfaced on the CI of an unrelated PR (#5234), where I could show the change was not causal — the file it touches is an empty translation unit in that build configuration. Filing it there would have buried it.
It is also the third distinct intermittent abort I have seen from this harness in the last few days. The other two were pj_throw_exception_ reporting no handler, and a PESQ media test failing after RTP socket bind() ... Address already in use. The last of those looks environmental; this one does not, which is why it seemed worth a separate report rather than being written off as flakiness.
PJSIP version
master, observed 2026-09-02.
Context
macOS runner, arm64, OpenSSL backend (PJ_SSL_SOCK_IMP : 1). Not Apple-TLS-specific — the assertion is in the pool allocator and the SSL socket implementation is incidental to it.
Suggested labels
type: bug, priority: low, and both component: pjlib and component: unit-tests — the assertion is in pjlib's allocator, but it is only observed through the test harness and the shuffled multi-worker configuration is what exposes it. I could not tell which side the fault is on, so I have not picked one.
Describe the bug
pjlib-testintermittently aborts on a pool-alignment assertion when run with--shuffleand more than one worker thread. Seen on the macOS CI runners:It aborts roughly 35 ms into the run, so before much has happened, and it is not reproducible on demand — the same commit passes the same job on other runs.
What the assertion means
The value cannot have come from the caller.
pj_pool_aligned_alloc()already screens it (pjlib/include/pj/pool_i.h:81):A caller passing
0— the ordinary "use the pool default" case — is accepted, and the value is then taken frompool->alignmentand passed topj_pool_alloc_from_block(), which asserts it is a power of two. So reaching that assertion meanspool->alignmentitself was0, or otherwise not a power of two, at the moment of the call.Two things follow, and they are separable:
1. The substituted value is never validated. The caller's alignment is screened on the line above; the pool's own is not. Whatever the underlying cause,
pool->alignmentbecoming invalid currently surfaces as an assertion three frames away from wherever the pool was actually damaged. Validating it at pool creation — so the field cannot hold a non-power-of-two — or screening it at the substitution would turn this into a diagnosable failure instead of a puzzle.2. Something is leaving
pool->alignmentinvalid. I have not diagnosed this and do not want to guess. Zero is what an uninitialised or zeroed pool header would hold, and the failure only appears under--shufflewith two workers, which is suggestive of a lifetime or concurrency fault — a pool used after release, or one being created concurrently with use — rather than a bad call site. That is a hypothesis, not a finding.Steps to reproduce
No reliable reproduction. It appears sporadically in CI under the existing shuffled multi-worker configuration:
The seed is printed on each run (
random seed=312above), so a failing ordering should be replayable if the harness accepts a fixed seed — that seems the cheapest first step toward a reproduction, and is why I have quoted it.Why I am filing it separately
It surfaced on the CI of an unrelated PR (#5234), where I could show the change was not causal — the file it touches is an empty translation unit in that build configuration. Filing it there would have buried it.
It is also the third distinct intermittent abort I have seen from this harness in the last few days. The other two were
pj_throw_exception_reporting no handler, and a PESQ media test failing afterRTP socket bind() ... Address already in use. The last of those looks environmental; this one does not, which is why it seemed worth a separate report rather than being written off as flakiness.PJSIP version
master, observed 2026-09-02.
Context
macOS runner, arm64, OpenSSL backend (
PJ_SSL_SOCK_IMP : 1). Not Apple-TLS-specific — the assertion is in the pool allocator and the SSL socket implementation is incidental to it.Suggested labels
type: bug,priority: low, and bothcomponent: pjlibandcomponent: unit-tests— the assertion is in pjlib's allocator, but it is only observed through the test harness and the shuffled multi-worker configuration is what exposes it. I could not tell which side the fault is on, so I have not picked one.