Repository navigation
Conversation
DMD perf check
Breakdown — compile hello.d
All measurements
c01a0b2 vs merge-base f302b81 · about these metrics |
a152334 to
e038f21
Compare
| enum pureBumpMalloc = cast(void* function(size_t) pure nothrow)&bumpMalloc; | ||
| enum pureBumpFree = cast(void function(void*p, size_t) pure nothrow)&bumpFree; |
There was a problem hiding this comment.
Could also fake it with pragma(mangle), which has support on older compilers.
dmd/druntime/src/core/memory.d
Line 1234 in 807e8a2
There was a problem hiding this comment.
Indeed, that seems to work across all versions. Updated.
There was a problem hiding this comment.
LDC is not happy with this change: Function type does not match previously declared function with the same mangled name. Restored.
12359e1 to
9502df6
Compare
|
Cherry-picking this on Update: Ahh, merge works. Nevermind this. |
|
Does this supersede #23914 ?
I'd say definitely yes! |
| enum smallAlignmentShift = 4; | ||
| enum smallAlignment = 1 << smallAlignmentShift; | ||
| enum maxSmallSize = 1024; | ||
| __gshared void*[maxSmallSize >> smallAlignmentShift] firstSmallFree; | ||
|
|
||
| enum largeAlignmentShift = 8; | ||
| enum largeAlignment = 1 << largeAlignmentShift; | ||
| enum maxLargeSize = 64 * 1024; | ||
| __gshared void*[maxLargeSize >> largeAlignmentShift] firstLargeFree; | ||
|
|
||
| enum hugeAlignment = 4096; | ||
| __gshared void* firstHugeFree; |
There was a problem hiding this comment.
Is this bucket strategy tuned for dmd?
There was a problem hiding this comment.
Is this bucket strategy tuned for dmd?
Not really, just my first guess. I suspect allocation sizes are mostly rather small, but I can add some statistics on bucket usage.
There was a problem hiding this comment.
Here's a summary of the bin usage of my phobos unittest benchmark:
| Size range | Mallocs | Frees | Approx. live memory | Interpretation |
|---|---|---|---|---|
0x01–0x10 |
8,587,199 | 3,549,229 | ~50 MB | Very high count |
0x11–0x20 |
11,973,412 | 188,068 | ~377 MB | Huge count + memory |
0x21–0x30 |
535,439 | 84,989 | ~21 MB | Moderate |
0x31–0x40 |
2,076,679 | 17,630 | ~131 MB | High count + memory |
0x41–0x50 |
7,573,155 | 36,854 | ~452 MB | Huge count + memory |
0x51–0x60 |
108,464 | 5,100 | ~6 MB | Low impact |
0x61–0x100 |
~1k–26k/bucket | ~0.5k–11k/bucket | <10 MB each | Low impact |
0x101–0x300 |
~700–6,000/bucket | ~400–1,100/bucket | <1 MB each | Negligible individually |
0x301–0x3000 |
hundreds–thousands | tens–thousands | <1 MB/bucket | Low |
0x3001–0xffff |
tens–thousands | generally similar | Generally small | Low |
>= 0x10000 |
89 | 57 | ~47 MB net reported | Few allocations, but expensive |
No, it replaces the C malloc calls by using the memory of the bump allocator, that can benefit additionally from #23914 when using larger pages. The measurements in #23914 (comment) were made on top of this PR. |
9199c40 to
a335419
Compare
For my usual testcase (all phobos unittests compiled with -o-), this reduces compilation time by about 9% (26.8 s -> 24.6 s) and reduces process memory by about 2.5 % (12212 MB -> 11943 MB).
Not sure if that's worth having to pass allocation sizes to mem.xfree and mem.xrealloc.
Let's see what the perf runner reports...