Repository navigation
Slow all to one copies #1232
Description
Activity
Could you please run with LEGATE_LOG_MAPPING=1 and LEGATE_SHOW_PROGRESS=1 and post the logging here? I want to confirm that Legate chose to use Reduction instances rather than "lifting" to read-write. Might be worth it to reduce the number of
--cpusper node, to reduce the size of the logging.It is using reduction instances:
[0 - 7f328d872740] 1.089779 {2}{legate.mapper}: MAP_TASK for fill(index_point=(6))<132> [0 - 7f328d872740] 1.089790 {2}{legate.mapper}: TARGET PROCS: 1d00000000000008 [0 - 7f328d872740] 1.089797 {2}{legate.mapper}: CHOSEN INSTANCES: [0 - 7f328d872740] 1.089805 {2}{legate.mapper}: Requirement[0](privilege=REDUCE,region=(4,(4,4),4),domain=<0,0,0>..<399,1023,1023>,fields=10000) [0 - 7f328d872740] 1.089811 {2}{legate.mapper}: Instance[4000000000000006](REDUCTION,region=(4,*,4),memory=1e00000000000000,domain=<0,0,0>..<399,1023,1023>,fields=10000,layout=SoA:XYZ)Would it make sense to use a collective instance here instead?
I'm pretty sure we do something like this in s3d using collective instances and dont have this problem. I might try doing something like this in regent and maybe playing with the mapper.
Would it make sense to use a collective instance here instead?
Yes, although Legate's mapper should in theory already be catching that case if it's possible. If you want to check if it's amenable to that and whether Legate's mapper is messing it up, you can capture lightweight Legion Spy logs and ask it to do the collective analysis.
Agreed, this case should be getting caught. Let me assign to @dinodeep to investigate. Deep, we want to check why
populate_input_collective_regionsis not catching this particular argument.Reacted by Deep PatelRunning legion_spy with --collective does not seem to say anything:
Reducing top-level index space shapes... |||||||||||||||||||||||||||||||||||||||||||||||||||| 100.0% Done Computing refinement points... Dim 3: |||||||||||||||||||||||||||||||||||||||||||||||||||| 100.0% Dim 1: |||||||||||||||||||||||||||||||||||||||||||||||||||| 100.0% Done Checking collective rendezvous... Legion Spy analysis complete. Exiting...So if Legion Spy isn't complaining that means it's probably not a case that can be fixed by a mapper change to use collective instances. More likely this is what @manopapad was suggesting earlier where Legate is promoting to READ_WRITE:
I want to confirm that Legate chose to use Reduction instances rather than "lifting" to read-write.
Then I'd like to see the full logging from these runs. Based on #1232 (comment) it appears that Reduction privileges are indeed in use.
Heres the rank 0 log where print_arr executes:
[0 - 7fbed731e240] 0.086755 {2}{legate.mapper}: Memories on rank 0: [0 - 7fbed731e240] 0.087477 {2}{legate.mapper}: 1e00000000000000 (SYSTEM_MEM): 34359738368 bytes [0 - 7fbed731e240] 0.087491 {2}{legate.mapper}: 1e00000000000001 (SYSTEM_MEM): 0 bytes [0 - 7fbed731e240] 0.087499 {2}{legate.mapper}: 1e00000000000002 (FILE_MEM): 0 bytes [0 - 7fbed731e240] 0.087506 {2}{legate.mapper}: Processors on rank 0: [0 - 7fbed731e240] 0.087514 {2}{legate.mapper}: 1d00000000000000 (UTIL_PROC) can see 1e00000000000000(bw=100) 1e00000000000001(bw=100) 1e00000000000002(bw=5) [0 - 7fbed731e240] 0.087521 {2}{legate.mapper}: 1d00000000000001 (UTIL_PROC) can see 1e00000000000000(bw=100) 1e00000000000001(bw=100) 1e00000000000002(bw=5) [0 - 7fbed731e240] 0.087529 {2}{legate.mapper}: 1d00000000000002 (LOC_PROC) can see 1e00000000000000(bw=100) 1e00000000000001(bw=100) 1e00000000000002(bw=5) [0 - 7fbed731e240] 0.087537 {2}{legate.mapper}: 1d00000000000003 (LOC_PROC) can see 1e00000000000000(bw=100) 1e00000000000001(bw=100) 1e00000000000002(bw=5) [0 - 7fbed731e240] 0.087544 {2}{legate.mapper}: 1d00000000000004 (LOC_PROC) can see 1e00000000000000(bw=100) 1e00000000000001(bw=100) 1e00000000000002(bw=5) [0 - 7fbed731e240] 0.087551 {2}{legate.mapper}: 1d00000000000005 (LOC_PROC) can see 1e00000000000000(bw=100) 1e00000000000001(bw=100) 1e00000000000002(bw=5) [0 - 7fb6747de740] 0.898452 {2}{legate.mapper}: SELECT_SHARDING_FUNCTOR for fill<24> [0 - 7fb6747de740] 0.898481 {2}{legate.mapper}: 0 <- (0) (1) (2) (3) [0 - 7fb6747de740] 0.898488 {2}{legate.mapper}: 1 <- (4) (5) (6) (7) [0 - 7fb6747de740] 0.898495 {2}{legate.mapper}: 2 <- (8) (9) (10) (11) [0 - 7fb6747de740] 0.898503 {2}{legate.mapper}: 3 <- (12) (13) (14) (15) [0 - 7fb6747de740] 0.898536 {2}{legate.mapper}: SELECT_TASK_OPTIONS for fill<24>: initial_proc=1d00000000000002 valid_instances=false [0 - 7fb6747de740] 0.899281 {2}{legate.mapper}: SELECT_SHARDING_FUNCTOR for print_arr<28> [0 - 7fb6747de740] 0.899291 {2}{legate.mapper}: 0 <- (0) [0 - 7fb6747de740] 0.899299 {2}{legate.mapper}: 1 <- [0 - 7fb6747de740] 0.899306 {2}{legate.mapper}: 2 <- [0 - 7fb6747de740] 0.899313 {2}{legate.mapper}: 3 <- [0 - 7fb6747de740] 0.899327 {2}{legate.mapper}: SELECT_TASK_OPTIONS for print_arr<28>: initial_proc=1d00000000000002 valid_instances=false [0 - 7fb6747de740] 0.931880 {2}{legate.mapper}: SLICE_TASK for fill<24> [0 - 7fb6747de740] 0.931896 {2}{legate.mapper}: <0>..<0> -> 1d00000000000002 [0 - 7fb6747de740] 0.931904 {2}{legate.mapper}: <1>..<1> -> 1d00000000000003 [0 - 7fb6747de740] 0.931911 {2}{legate.mapper}: <2>..<2> -> 1d00000000000004 [0 - 7fb6747de740] 0.931917 {2}{legate.mapper}: <3>..<3> -> 1d00000000000005 [0 - 7fb6747de740] 0.957687 {2}{legate.mapper}: MAP_TASK for fill(index_point=(0))<76> [0 - 7fb6747de740] 0.957699 {2}{legate.mapper}: TARGET PROCS: 1d00000000000002 [0 - 7fb6747de740] 0.957706 {2}{legate.mapper}: CHOSEN INSTANCES: [0 - 7fb6747de740] 0.957713 {2}{legate.mapper}: Requirement[0](privilege=REDUCE,region=(4,(4,4),4),domain=<0,0,0>..<399,1023,1023>,fields=10000) [0 - 7fb6747de740] 0.957721 {2}{legate.mapper}: Instance[4000000000000000](REDUCTION,region=(4,*,4),memory=1e00000000000000,domain=<0,0,0>..<399,1023,1023>,fields=10000,layout=SoA:XYZ) [0 - 7fb6747de740] 0.957728 {2}{legate.mapper}: MEMORY POOLS: [0 - 7fb6747de740] 0.957736 {2}{legate.mapper}: Memory 1e00000000000000 of kind SYSTEM_MEM: BOUNDED_POOL(size=0,alignment=16) [0 - 7fb6747d2740] 0.961150 {2}{legate.mapper}: MAP_TASK for fill(index_point=(1))<80> [0 - 7fb6747d2740] 0.961164 {2}{legate.mapper}: TARGET PROCS: 1d00000000000003 [0 - 7fb6747d2740] 0.961172 {2}{legate.mapper}: CHOSEN INSTANCES: [0 - 7fb6747d2740] 0.961180 {2}{legate.mapper}: Requirement[0](privilege=REDUCE,region=(4,(4,4),4),domain=<0,0,0>..<399,1023,1023>,fields=10000) [0 - 7fb6747d2740] 0.961187 {2}{legate.mapper}: Instance[4000000000000001](REDUCTION,region=(4,*,4),memory=1e00000000000000,domain=<0,0,0>..<399,1023,1023>,fields=10000,layout=SoA:XYZ) [0 - 7fb6747d2740] 0.961196 {2}{legate.mapper}: MEMORY POOLS: [0 - 7fb6747d2740] 0.961204 {2}{legate.mapper}: Memory 1e00000000000000 of kind SYSTEM_MEM: BOUNDED_POOL(size=0,alignment=16) [0 - 7fb6747de740] 0.962052 {2}{legate.mapper}: MAP_TASK for fill(index_point=(2))<84> [0 - 7fb6747de740] 0.962069 {2}{legate.mapper}: TARGET PROCS: 1d00000000000004 [0 - 7fb6747de740] 0.962077 {2}{legate.mapper}: CHOSEN INSTANCES: [0 - 7fb6747de740] 0.962085 {2}{legate.mapper}: Requirement[0](privilege=REDUCE,region=(4,(4,4),4),domain=<0,0,0>..<399,1023,1023>,fields=10000) [0 - 7fb6747de740] 0.962093 {2}{legate.mapper}: Instance[4000000000000002](REDUCTION,region=(4,*,4),memory=1e00000000000000,domain=<0,0,0>..<399,1023,1023>,fields=10000,layout=SoA:XYZ) [0 - 7fb6747de740] 0.962100 {2}{legate.mapper}: MEMORY POOLS: [0 - 7fb6747de740] 0.962107 {2}{legate.mapper}: Memory 1e00000000000000 of kind SYSTEM_MEM: BOUNDED_POOL(size=0,alignment=16) [0 - 7fb6747de740] 0.962243 {2}{legate.mapper}: MAP_TASK for fill(index_point=(3))<88> [0 - 7fb6747de740] 0.962256 {2}{legate.mapper}: TARGET PROCS: 1d00000000000005 [0 - 7fb6747de740] 0.962265 {2}{legate.mapper}: CHOSEN INSTANCES: [0 - 7fb6747de740] 0.962272 {2}{legate.mapper}: Requirement[0](privilege=REDUCE,region=(4,(4,4),4),domain=<0,0,0>..<399,1023,1023>,fields=10000) [0 - 7fb6747de740] 0.962280 {2}{legate.mapper}: Instance[4000000000000003](REDUCTION,region=(4,*,4),memory=1e00000000000000,domain=<0,0,0>..<399,1023,1023>,fields=10000,layout=SoA:XYZ) [0 - 7fb6747de740] 0.962287 {2}{legate.mapper}: MEMORY POOLS: [0 - 7fb6747de740] 0.962294 {2}{legate.mapper}: Memory 1e00000000000000 of kind SYSTEM_MEM: BOUNDED_POOL(size=0,alignment=16) [0 - 7fb6747de740] 0.987742 {2}{legate.mapper}: SLICE_TASK for print_arr<28> [0 - 7fb6747de740] 0.987768 {2}{legate.mapper}: <0>..<0> -> 1d00000000000002 [0 - 7fb6747de740] 0.987961 {2}{legate.mapper}: MAP_TASK for print_arr(index_point=(0))<96> [0 - 7fb6747de740] 0.987975 {2}{legate.mapper}: TARGET PROCS: 1d00000000000002 [0 - 7fb6747de740] 0.987982 {2}{legate.mapper}: CHOSEN INSTANCES: [0 - 7fb6747de740] 0.987990 {2}{legate.mapper}: Requirement[0](privilege=READ_ONLY,region=(4,(4,4),4),domain=<0,0,0>..<399,1023,1023>,fields=10000) [0 - 7fb6747de740] 0.987997 {2}{legate.mapper}: Instance[4000000000000004](region=(4,*,4),memory=1e00000000000000,domain=<0,0,0>..<399,1023,1023>,fields=10000,layout=SoA:XYZ) [0 - 7fb6747de740] 0.988005 {2}{legate.mapper}: MEMORY POOLS: [0 - 7fb6747de740] 0.988012 {2}{legate.mapper}: Memory 1e00000000000000 of kind SYSTEM_MEM: BOUNDED_POOL(size=0,alignment=16) [0 - 7fb6747a2740] 2.439099 {3}{legate}: fill CPU task [], pt = (3), proc = 1d00000000000005 [0 - 7fb6747ba740] 2.440736 {3}{legate}: fill CPU task [], pt = (1), proc = 1d00000000000003 [0 - 7fb6747c6740] 2.459115 {3}{legate}: fill CPU task [], pt = (0), proc = 1d00000000000002 [0 - 7fb6747ae740] 2.493321 {3}{legate}: fill CPU task [], pt = (2), proc = 1d00000000000004 [0 - 7fb6747c6740] 56.648769 {3}{legate}: print_arr CPU task [], pt = (0), proc = 1d00000000000002Maybe a new read_only instance is getting created?
Everything appears to be correct. AFAICT Legate should be enabling collective instances, and Legion spy should be complaining that we're not. The only "suspicious" thing is that you're doing a manual task launch, but that's still an index launch (as we can confirm from the mapper logs). So I'm inclined to wait for @dinodeep to try enabling collective instances in
select_task_optionsfor this case, and presumably Legion will inform us what is stopping this.If Legion Spy is not complaining, I don't think that enabling collective instances in
select_task_optionsis going to do anything. You're welcome to try though.Will Legion print a warning though, that we requested collective instances when we shouldn't have? And ideally why collective use was impossible?
10 remaining items
Sorry for the late reply. Yes, currently the latest cupynumeric packages don't support these new Legate changes because the Legate SHA has not yet been updated to use it. You can build from source to get the changes, or we have an internal cupynumeric PR that will be updating the Legate SHA, and I can ping here when it is merged in
@syamajala you should be able to test the solution using the latest nightlies.
@syamajala (Seshu), please confirm if you were able to validate if the solution works. Thanks
It does look like we're using collective instances now and I see fewer large copies instead of many small ones, but the copies are still the bottleneck:
https://legion.stanford.edu/prof-viewer/?url=https://sapling2.stanford.edu/~seshu/legion_prof_legate2/Here's my read based on the profile:
It makes sense that CPU tasks are taking a while to complete. We have 8 CPU processors, but they are executing python code, so they will get serialized due to the GIL.
That said, I don't understand why the 8 CPU tasks on each node are finishing together, maybe @lightsighter has a suggestion?
Each of the tasks is using a separate reduction instance, this appears to be a choice that the Legate mapper is making https://github.com/nv-legate/legate/blob/55b62eccec288e126d0875184b3ab0b06a66f62c/src/cpp/legate/mapping/detail/base_mapper.cc#L858. @ipdemes do you recall why we chose not to cache reduction instances for CPU tasks?
@syamajala you could try to enable the above option in the mapper for CPU tasks too (assuming you're set up for a source build), or we can make a one-off build for you. The execution of CPU tasks will be serialized, but you already have this problem due to using Python tasks.
The copies on each node are one-by-one filling in the "final" instance based on each task's intermediate reduction results. E.g. on node 0, 0x4000000000000000 is that instance. Legion could have possibly chosen to do a "node local" tree reduction, but if you're bottlenecked on local reduction copy bandwidth then it wouldn't make a difference.
However, it appears that intra-csize reduction copies have effective bandwidth around 200MB/s. This is way lower than remote reduction copies, that are closer to 5GB/s. @eddy16112 @apryakhin do you know why this might be?
The tasks on nodes 1,2,3 are finishing way late, and that's a bit contributor to the latency. I don't know why that would be, maybe an nsys profile could tell us. These are running on separate nodes right?
I can try building from source again and play with the mapper. I had a lot of issues before due to the Realm split, but I assume those have all been fixed at this point, so I'll try again.
That was all running on 1 node.
I had a lot of issues before due to the Realm split, but I assume those have all been fixed at this point, so I'll try again.
Yes, there is now a
--with-realm-src-diralso you can use.That said, I don't understand why the 8 CPU tasks on each node are finishing together, maybe @lightsighter has a suggestion?
Usually I would say that is the result of some kind of synchronization between them, e.g. they are all doing a collective of some kind or synchronizing on a barrier. It would have to be something that is not a Realm/Legion primitive though because if it were then they would actually wait on an event and that would show up in the profile as the tasks waiting and being de-scheduled. Here it looks like they are all running concurrently until the point where they all finish.
- added a commit that references this issue
on Dec 17, 2025 @syamajala is this still an issue? If not, can we go ahead and close this and revisit as needed.
This has been inactive for a while, closing. @syamajala please reopen if you still need this.
Reacted by Seshu Yamajala

Is there some way to improve the performance of this example:
There is a profile here: https://legion.stanford.edu/prof-viewer/?url=https://sapling2.stanford.edu/~seshu/legion_prof_legate/