Skip to content
This repository was archived by the owner on Oct 5, 2026. It is now read-only.
This repository was archived by the owner on Oct 5, 2026. It is now read-only.

Slow all to one copies #1232

Description

@syamajala

Is there some way to improve the performance of this example:

import numpy as np
import cupynumeric as cpy
from legate.core.task import task, InputStore, OutputStore, ReductionStore, ADD
from legate.core import(
    VariantCode,
    broadcast,
    align,
    dimension,
    constant,
    get_legate_runtime,
    LegateDataInterface,
    LogicalStore,
    get_machine,
    TaskTarget,
    Machine
    )

def get_store(obj: LegateDataInterface) -> LogicalStore:
    iface = obj.__legate_data_interface__
    assert iface["version"] == 1
    data = iface["data"]
    # There should only be one field
    assert len(data) == 1
    field = next(iter(data))
    assert not field.nullable
    column = data[field]
    assert not column.nullable
    return column.data

@task(
    variants=(VariantCode.CPU,)
)
def fill(arr_store : ReductionStore[ADD]):
    arr = np.asarray(arr_store)
    arr += np.ones(arr.shape)

@task(
    variants=(VariantCode.CPU,)
)
def print_arr(arr_store : InputStore):
    arr = np.asarray(arr_store)
    print(arr[0])

machine = get_machine()
cpus = machine.only(TaskTarget.CPU).count()
print("CPUS:", cpus)

arr = cpy.zeros((400, 1024, 1024))

runtime = get_legate_runtime()
library = fill.library
fill_task = runtime.create_manual_task(library, fill.task_id, (cpus,))
fill_task.add_reduction(get_store(arr), ADD)
fill_task.execute()

library = print_arr.library
print_arr_task = runtime.create_manual_task(library, print_arr.task_id, (1,))
print_arr_task.add_input(get_store(arr))
print_arr_task.execute()

There is a profile here: https://legion.stanford.edu/prof-viewer/?url=https://sapling2.stanford.edu/~seshu/legion_prof_legate/

Activity

  1. manopapad commented on Sep 10, 2025

    @manopapad
    Contributor

    Could you please run with LEGATE_LOG_MAPPING=1 and LEGATE_SHOW_PROGRESS=1 and post the logging here? I want to confirm that Legate chose to use Reduction instances rather than "lifting" to read-write. Might be worth it to reduce the number of --cpus per node, to reduce the size of the logging.

  2. added theissue type on Sep 10, 2025
  3. syamajala commented on Sep 10, 2025

    @syamajala
    Author

    It is using reduction instances:

    [0 - 7f328d872740]    1.089779 {2}{legate.mapper}: MAP_TASK for fill(index_point=(6))<132>
    [0 - 7f328d872740]    1.089790 {2}{legate.mapper}:   TARGET PROCS: 1d00000000000008
    [0 - 7f328d872740]    1.089797 {2}{legate.mapper}:   CHOSEN INSTANCES:
    [0 - 7f328d872740]    1.089805 {2}{legate.mapper}:     Requirement[0](privilege=REDUCE,region=(4,(4,4),4),domain=<0,0,0>..<399,1023,1023>,fields=10000)
    [0 - 7f328d872740]    1.089811 {2}{legate.mapper}:       Instance[4000000000000006](REDUCTION,region=(4,*,4),memory=1e00000000000000,domain=<0,0,0>..<399,1023,1023>,fields=10000,layout=SoA:XYZ)
    

    Would it make sense to use a collective instance here instead?

  4. syamajala commented on Sep 10, 2025

    @syamajala
    Author

    I'm pretty sure we do something like this in s3d using collective instances and dont have this problem. I might try doing something like this in regent and maybe playing with the mapper.

  5. lightsighter commented on Sep 10, 2025

    @lightsighter
    Contributor

    Would it make sense to use a collective instance here instead?

    Yes, although Legate's mapper should in theory already be catching that case if it's possible. If you want to check if it's amenable to that and whether Legate's mapper is messing it up, you can capture lightweight Legion Spy logs and ask it to do the collective analysis.

  6. manopapad commented on Sep 10, 2025

    @manopapad
    Contributor

    Agreed, this case should be getting caught. Let me assign to @dinodeep to investigate. Deep, we want to check why populate_input_collective_regions is not catching this particular argument.

  7. syamajala commented on Sep 10, 2025

    @syamajala
    Author

    Running legion_spy with --collective does not seem to say anything:

    Reducing top-level index space shapes...
     |||||||||||||||||||||||||||||||||||||||||||||||||||| 100.0%                                                          
    Done
    Computing refinement points...
    Dim 3: |||||||||||||||||||||||||||||||||||||||||||||||||||| 100.0%                                                    
    Dim 1: |||||||||||||||||||||||||||||||||||||||||||||||||||| 100.0%                                                    
    Done
    Checking collective rendezvous...
    Legion Spy analysis complete.  Exiting...
    
  8. lightsighter commented on Sep 10, 2025

    @lightsighter
    Contributor

    So if Legion Spy isn't complaining that means it's probably not a case that can be fixed by a mapper change to use collective instances. More likely this is what @manopapad was suggesting earlier where Legate is promoting to READ_WRITE:

    I want to confirm that Legate chose to use Reduction instances rather than "lifting" to read-write.

  9. manopapad commented on Sep 10, 2025

    @manopapad
    Contributor

    Then I'd like to see the full logging from these runs. Based on #1232 (comment) it appears that Reduction privileges are indeed in use.

  10. syamajala commented on Sep 10, 2025

    @syamajala
    Author

    Heres the rank 0 log where print_arr executes:

    [0 - 7fbed731e240]    0.086755 {2}{legate.mapper}: Memories on rank 0:
    [0 - 7fbed731e240]    0.087477 {2}{legate.mapper}:   1e00000000000000 (SYSTEM_MEM): 34359738368 bytes
    [0 - 7fbed731e240]    0.087491 {2}{legate.mapper}:   1e00000000000001 (SYSTEM_MEM): 0 bytes
    [0 - 7fbed731e240]    0.087499 {2}{legate.mapper}:   1e00000000000002 (FILE_MEM): 0 bytes
    [0 - 7fbed731e240]    0.087506 {2}{legate.mapper}: Processors on rank 0:
    [0 - 7fbed731e240]    0.087514 {2}{legate.mapper}:   1d00000000000000 (UTIL_PROC) can see 1e00000000000000(bw=100) 1e00000000000001(bw=100) 1e00000000000002(bw=5)
    [0 - 7fbed731e240]    0.087521 {2}{legate.mapper}:   1d00000000000001 (UTIL_PROC) can see 1e00000000000000(bw=100) 1e00000000000001(bw=100) 1e00000000000002(bw=5)
    [0 - 7fbed731e240]    0.087529 {2}{legate.mapper}:   1d00000000000002 (LOC_PROC) can see 1e00000000000000(bw=100) 1e00000000000001(bw=100) 1e00000000000002(bw=5)
    [0 - 7fbed731e240]    0.087537 {2}{legate.mapper}:   1d00000000000003 (LOC_PROC) can see 1e00000000000000(bw=100) 1e00000000000001(bw=100) 1e00000000000002(bw=5)
    [0 - 7fbed731e240]    0.087544 {2}{legate.mapper}:   1d00000000000004 (LOC_PROC) can see 1e00000000000000(bw=100) 1e00000000000001(bw=100) 1e00000000000002(bw=5)
    [0 - 7fbed731e240]    0.087551 {2}{legate.mapper}:   1d00000000000005 (LOC_PROC) can see 1e00000000000000(bw=100) 1e00000000000001(bw=100) 1e00000000000002(bw=5)
    [0 - 7fb6747de740]    0.898452 {2}{legate.mapper}: SELECT_SHARDING_FUNCTOR for fill<24>
    [0 - 7fb6747de740]    0.898481 {2}{legate.mapper}:   0 <- (0) (1) (2) (3)
    [0 - 7fb6747de740]    0.898488 {2}{legate.mapper}:   1 <- (4) (5) (6) (7)
    [0 - 7fb6747de740]    0.898495 {2}{legate.mapper}:   2 <- (8) (9) (10) (11)
    [0 - 7fb6747de740]    0.898503 {2}{legate.mapper}:   3 <- (12) (13) (14) (15)
    [0 - 7fb6747de740]    0.898536 {2}{legate.mapper}: SELECT_TASK_OPTIONS for fill<24>: initial_proc=1d00000000000002 valid_instances=false
    [0 - 7fb6747de740]    0.899281 {2}{legate.mapper}: SELECT_SHARDING_FUNCTOR for print_arr<28>
    [0 - 7fb6747de740]    0.899291 {2}{legate.mapper}:   0 <- (0)
    [0 - 7fb6747de740]    0.899299 {2}{legate.mapper}:   1 <-
    [0 - 7fb6747de740]    0.899306 {2}{legate.mapper}:   2 <-
    [0 - 7fb6747de740]    0.899313 {2}{legate.mapper}:   3 <-
    [0 - 7fb6747de740]    0.899327 {2}{legate.mapper}: SELECT_TASK_OPTIONS for print_arr<28>: initial_proc=1d00000000000002 valid_instances=false
    [0 - 7fb6747de740]    0.931880 {2}{legate.mapper}: SLICE_TASK for fill<24>
    [0 - 7fb6747de740]    0.931896 {2}{legate.mapper}:   <0>..<0> -> 1d00000000000002
    [0 - 7fb6747de740]    0.931904 {2}{legate.mapper}:   <1>..<1> -> 1d00000000000003
    [0 - 7fb6747de740]    0.931911 {2}{legate.mapper}:   <2>..<2> -> 1d00000000000004
    [0 - 7fb6747de740]    0.931917 {2}{legate.mapper}:   <3>..<3> -> 1d00000000000005
    [0 - 7fb6747de740]    0.957687 {2}{legate.mapper}: MAP_TASK for fill(index_point=(0))<76>
    [0 - 7fb6747de740]    0.957699 {2}{legate.mapper}:   TARGET PROCS: 1d00000000000002
    [0 - 7fb6747de740]    0.957706 {2}{legate.mapper}:   CHOSEN INSTANCES:
    [0 - 7fb6747de740]    0.957713 {2}{legate.mapper}:     Requirement[0](privilege=REDUCE,region=(4,(4,4),4),domain=<0,0,0>..<399,1023,1023>,fields=10000)
    [0 - 7fb6747de740]    0.957721 {2}{legate.mapper}:       Instance[4000000000000000](REDUCTION,region=(4,*,4),memory=1e00000000000000,domain=<0,0,0>..<399,1023,1023>,fields=10000,layout=SoA:XYZ)
    [0 - 7fb6747de740]    0.957728 {2}{legate.mapper}:   MEMORY POOLS:
    [0 - 7fb6747de740]    0.957736 {2}{legate.mapper}:     Memory 1e00000000000000 of kind SYSTEM_MEM: BOUNDED_POOL(size=0,alignment=16)
    [0 - 7fb6747d2740]    0.961150 {2}{legate.mapper}: MAP_TASK for fill(index_point=(1))<80>
    [0 - 7fb6747d2740]    0.961164 {2}{legate.mapper}:   TARGET PROCS: 1d00000000000003
    [0 - 7fb6747d2740]    0.961172 {2}{legate.mapper}:   CHOSEN INSTANCES:
    [0 - 7fb6747d2740]    0.961180 {2}{legate.mapper}:     Requirement[0](privilege=REDUCE,region=(4,(4,4),4),domain=<0,0,0>..<399,1023,1023>,fields=10000)
    [0 - 7fb6747d2740]    0.961187 {2}{legate.mapper}:       Instance[4000000000000001](REDUCTION,region=(4,*,4),memory=1e00000000000000,domain=<0,0,0>..<399,1023,1023>,fields=10000,layout=SoA:XYZ)
    [0 - 7fb6747d2740]    0.961196 {2}{legate.mapper}:   MEMORY POOLS:
    [0 - 7fb6747d2740]    0.961204 {2}{legate.mapper}:     Memory 1e00000000000000 of kind SYSTEM_MEM: BOUNDED_POOL(size=0,alignment=16)
    [0 - 7fb6747de740]    0.962052 {2}{legate.mapper}: MAP_TASK for fill(index_point=(2))<84>
    [0 - 7fb6747de740]    0.962069 {2}{legate.mapper}:   TARGET PROCS: 1d00000000000004
    [0 - 7fb6747de740]    0.962077 {2}{legate.mapper}:   CHOSEN INSTANCES:
    [0 - 7fb6747de740]    0.962085 {2}{legate.mapper}:     Requirement[0](privilege=REDUCE,region=(4,(4,4),4),domain=<0,0,0>..<399,1023,1023>,fields=10000)
    [0 - 7fb6747de740]    0.962093 {2}{legate.mapper}:       Instance[4000000000000002](REDUCTION,region=(4,*,4),memory=1e00000000000000,domain=<0,0,0>..<399,1023,1023>,fields=10000,layout=SoA:XYZ)
    [0 - 7fb6747de740]    0.962100 {2}{legate.mapper}:   MEMORY POOLS:
    [0 - 7fb6747de740]    0.962107 {2}{legate.mapper}:     Memory 1e00000000000000 of kind SYSTEM_MEM: BOUNDED_POOL(size=0,alignment=16)
    [0 - 7fb6747de740]    0.962243 {2}{legate.mapper}: MAP_TASK for fill(index_point=(3))<88>
    [0 - 7fb6747de740]    0.962256 {2}{legate.mapper}:   TARGET PROCS: 1d00000000000005
    [0 - 7fb6747de740]    0.962265 {2}{legate.mapper}:   CHOSEN INSTANCES:
    [0 - 7fb6747de740]    0.962272 {2}{legate.mapper}:     Requirement[0](privilege=REDUCE,region=(4,(4,4),4),domain=<0,0,0>..<399,1023,1023>,fields=10000)
    [0 - 7fb6747de740]    0.962280 {2}{legate.mapper}:       Instance[4000000000000003](REDUCTION,region=(4,*,4),memory=1e00000000000000,domain=<0,0,0>..<399,1023,1023>,fields=10000,layout=SoA:XYZ)
    [0 - 7fb6747de740]    0.962287 {2}{legate.mapper}:   MEMORY POOLS:
    [0 - 7fb6747de740]    0.962294 {2}{legate.mapper}:     Memory 1e00000000000000 of kind SYSTEM_MEM: BOUNDED_POOL(size=0,alignment=16)
    [0 - 7fb6747de740]    0.987742 {2}{legate.mapper}: SLICE_TASK for print_arr<28>
    [0 - 7fb6747de740]    0.987768 {2}{legate.mapper}:   <0>..<0> -> 1d00000000000002
    [0 - 7fb6747de740]    0.987961 {2}{legate.mapper}: MAP_TASK for print_arr(index_point=(0))<96>
    [0 - 7fb6747de740]    0.987975 {2}{legate.mapper}:   TARGET PROCS: 1d00000000000002
    [0 - 7fb6747de740]    0.987982 {2}{legate.mapper}:   CHOSEN INSTANCES:
    [0 - 7fb6747de740]    0.987990 {2}{legate.mapper}:     Requirement[0](privilege=READ_ONLY,region=(4,(4,4),4),domain=<0,0,0>..<399,1023,1023>,fields=10000)
    [0 - 7fb6747de740]    0.987997 {2}{legate.mapper}:       Instance[4000000000000004](region=(4,*,4),memory=1e00000000000000,domain=<0,0,0>..<399,1023,1023>,fields=10000,layout=SoA:XYZ)
    [0 - 7fb6747de740]    0.988005 {2}{legate.mapper}:   MEMORY POOLS:
    [0 - 7fb6747de740]    0.988012 {2}{legate.mapper}:     Memory 1e00000000000000 of kind SYSTEM_MEM: BOUNDED_POOL(size=0,alignment=16)
    [0 - 7fb6747a2740]    2.439099 {3}{legate}: fill CPU task [], pt = (3), proc = 1d00000000000005
    [0 - 7fb6747ba740]    2.440736 {3}{legate}: fill CPU task [], pt = (1), proc = 1d00000000000003
    [0 - 7fb6747c6740]    2.459115 {3}{legate}: fill CPU task [], pt = (0), proc = 1d00000000000002
    [0 - 7fb6747ae740]    2.493321 {3}{legate}: fill CPU task [], pt = (2), proc = 1d00000000000004
    [0 - 7fb6747c6740]   56.648769 {3}{legate}: print_arr CPU task [], pt = (0), proc = 1d00000000000002
    
  11. syamajala commented on Sep 10, 2025

    @syamajala
    Author

    Maybe a new read_only instance is getting created?

  12. manopapad commented on Sep 15, 2025

    @manopapad
    Contributor

    Everything appears to be correct. AFAICT Legate should be enabling collective instances, and Legion spy should be complaining that we're not. The only "suspicious" thing is that you're doing a manual task launch, but that's still an index launch (as we can confirm from the mapper logs). So I'm inclined to wait for @dinodeep to try enabling collective instances in select_task_options for this case, and presumably Legion will inform us what is stopping this.

  13. lightsighter commented on Sep 15, 2025

    @lightsighter
    Contributor

    If Legion Spy is not complaining, I don't think that enabling collective instances in select_task_options is going to do anything. You're welcome to try though.

  14. manopapad commented on Sep 15, 2025

    @manopapad
    Contributor

    Will Legion print a warning though, that we requested collective instances when we shouldn't have? And ideally why collective use was impossible?

  15. 10 remaining items

  16. dinodeep commented on Sep 29, 2025

    @dinodeep
    Contributor

    Sorry for the late reply. Yes, currently the latest cupynumeric packages don't support these new Legate changes because the Legate SHA has not yet been updated to use it. You can build from source to get the changes, or we have an internal cupynumeric PR that will be updating the Legate SHA, and I can ping here when it is merged in

  17. manopapad commented on Oct 8, 2025

    @manopapad
    Contributor

    @syamajala you should be able to test the solution using the latest nightlies.

  18. sbuneti commented on Oct 13, 2025

    @sbuneti

    @syamajala (Seshu), please confirm if you were able to validate if the solution works. Thanks

  19. syamajala commented on Oct 14, 2025

    @syamajala
    Author

    It does look like we're using collective instances now and I see fewer large copies instead of many small ones, but the copies are still the bottleneck:
    https://legion.stanford.edu/prof-viewer/?url=https://sapling2.stanford.edu/~seshu/legion_prof_legate2/

  20. syamajala commented on Oct 14, 2025

    @syamajala
    Author

    Im not sure I understand why some of these copies are happening? It seems like we should just be able to reuse the instances without making copies.

    Image
  21. manopapad commented on Oct 15, 2025

    @manopapad
    Contributor

    Here's my read based on the profile:

    It makes sense that CPU tasks are taking a while to complete. We have 8 CPU processors, but they are executing python code, so they will get serialized due to the GIL.

    That said, I don't understand why the 8 CPU tasks on each node are finishing together, maybe @lightsighter has a suggestion?

    Image

    Each of the tasks is using a separate reduction instance, this appears to be a choice that the Legate mapper is making https://github.com/nv-legate/legate/blob/55b62eccec288e126d0875184b3ab0b06a66f62c/src/cpp/legate/mapping/detail/base_mapper.cc#L858. @ipdemes do you recall why we chose not to cache reduction instances for CPU tasks?

    @syamajala you could try to enable the above option in the mapper for CPU tasks too (assuming you're set up for a source build), or we can make a one-off build for you. The execution of CPU tasks will be serialized, but you already have this problem due to using Python tasks.

    The copies on each node are one-by-one filling in the "final" instance based on each task's intermediate reduction results. E.g. on node 0, 0x4000000000000000 is that instance. Legion could have possibly chosen to do a "node local" tree reduction, but if you're bottlenecked on local reduction copy bandwidth then it wouldn't make a difference.

    Image

    However, it appears that intra-csize reduction copies have effective bandwidth around 200MB/s. This is way lower than remote reduction copies, that are closer to 5GB/s. @eddy16112 @apryakhin do you know why this might be?

    The tasks on nodes 1,2,3 are finishing way late, and that's a bit contributor to the latency. I don't know why that would be, maybe an nsys profile could tell us. These are running on separate nodes right?

  22. syamajala commented on Oct 15, 2025

    @syamajala
    Author

    I can try building from source again and play with the mapper. I had a lot of issues before due to the Realm split, but I assume those have all been fixed at this point, so I'll try again.

    That was all running on 1 node.

  23. manopapad commented on Oct 15, 2025

    @manopapad
    Contributor

    I had a lot of issues before due to the Realm split, but I assume those have all been fixed at this point, so I'll try again.

    Yes, there is now a --with-realm-src-dir also you can use.

  24. lightsighter commented on Oct 20, 2025

    @lightsighter
    Contributor

    That said, I don't understand why the 8 CPU tasks on each node are finishing together, maybe @lightsighter has a suggestion?

    Usually I would say that is the result of some kind of synchronization between them, e.g. they are all doing a collective of some kind or synchronizing on a barrier. It would have to be something that is not a Realm/Legion primitive though because if it were then they would actually wait on an event and that would show up in the profile as the tasks waiting and being de-scheduled. Here it looks like they are all running concurrently until the point where they all finish.

  25. sbuneti commented on Apr 13, 2026

    @sbuneti

    @syamajala is this still an issue? If not, can we go ahead and close this and revisit as needed.

  26. manopapad commented on Aug 6, 2026

    @manopapad
    Contributor

    This has been inactive for a while, closing. @syamajala please reopen if you still need this.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions