openmpi

Автор	SHA1	Сообщение	Дата
Nathan Hjelm	a06e491c2c	ob1: large buffered sends were broken by the ob1 optimizations. fix them The problem was caused by the static request optimization. The buffered send case is much like the isend case in that the request structure may be needed after MPI_Bsend completes. Fix this case by calling isend and freeing the resulting request. cmr=v1.7.5:ticket=trac:4149 This commit was SVN r30601. The following Trac tickets were found above: Ticket 4149 --> https://svn.open-mpi.org/trac/ompi/ticket/4149	2014-02-07 00:12:36 +00:00
Jeff Squyres	7d2c4cb468	There's a few ml-related bugs outstanding, and Nathan is looking into them, but it's going to take a little time (at least one day). So Nathan says it's ok to .ompi_ignore coll ml until he's able to fix it. This commit was SVN r30600.	2014-02-06 23:51:03 +00:00
Nathan Hjelm	3902cf66f1	ob1: OBJ_CONSTRUCT the convertor in the send_inline optimization. This change does not appear to increase the small message latency of ping-pong benchmarks and fixes an issue found by our ibm datatype tests. Fixes trac:4232 cmr=v1.7.5:ticket=trac:4149 This commit was SVN r30598. The following Trac tickets were found above: Ticket 4149 --> https://svn.open-mpi.org/trac/ompi/ticket/4149 Ticket 4232 --> https://svn.open-mpi.org/trac/ompi/ticket/4232	2014-02-06 21:27:42 +00:00
Nathan Hjelm	a41cb1f086	Remove duplicate definition of xpmem_apid_t cmr=v1.7.5:ticket=trac:4216 This commit was SVN r30589. The following Trac tickets were found above: Ticket 4216 --> https://svn.open-mpi.org/trac/ompi/ticket/4216	2014-02-06 20:38:20 +00:00
George Bosilca	6ee06b7fda	No exit down into a BTL. This commit was SVN r30566.	2014-02-05 15:04:01 +00:00
Ralph Castain	1326ed704f	Per the RFC discussed here: http://www.open-mpi.org/community/lists/devel/2014/01/13789.php add support for async modex when requested. cmr=v1.7.5:reviewer=jsquyres:subject=Add async modex support This commit was SVN r30565.	2014-02-05 14:39:27 +00:00
Joshua Ladd	1dbd8688db	This fixes a long standing bug in the OpenIB BTL's MCA param intialization.Only caught if BTL_OPENIB_FAILOVER_ENABLED. Thanks to Jeff for spotting. This should be added to: cmr=v1.7.4:reviewer=jsquyres cmr=v1.6.6 This commit was SVN r30558.	2014-02-04 20:01:39 +00:00
Nathan Hjelm	12f0bf9488	basesmuma: missed a couple of MB references cmr=v1.7.5:ticket=trac:4158 This commit was SVN r30538. The following Trac tickets were found above: Ticket 4158 --> https://svn.open-mpi.org/trac/ompi/ticket/4158	2014-02-03 18:19:53 +00:00
Nathan Hjelm	84320f3815	btl/vader: fix compilation with SGI xpmem and add some debugging to component_init. cmr=v1.7.5:ticker=#4053 This commit was SVN r30535.	2014-02-03 17:42:40 +00:00
Nathan Hjelm	64321acc22	basesmuma: do not call MB directly opal does not always define MB. It is recommended that opal_atomic_[rw]mb is called instead. We will need to address the cases where these functions are no-ops on weak-memory ordered cpus. cmr=v1.7.5:ticket=trac:4158 This commit was SVN r30534. The following Trac tickets were found above: Ticket 4158 --> https://svn.open-mpi.org/trac/ompi/ticket/4158	2014-02-03 17:01:57 +00:00
Nathan Hjelm	c2b061cc84	basesmuma: clean up code Several changes are contained in this commit: - Clean up tabs and trailing whitespaces - Use consistent indentation in changed files - Remove unused code. None of the removed code will ever have been used in a trunk build. - Clean up the smcm code quite a bit - Do not fflush stderr and use opal_output instead of fprintf. These changes have been tested on Cray XE-6 and PSM systems. cmr=v1.7.5:ticket=trac:4158 This commit was SVN r30533. The following Trac tickets were found above: Ticket 4158 --> https://svn.open-mpi.org/trac/ompi/ticket/4158	2014-02-03 17:01:46 +00:00
Christoph Niethammer	4f23d8214c	Fixed incorrect calculation of reallocated memory in mca_bml_r2_del_btl. This commit was SVN r30529.	2014-02-03 08:43:59 +00:00
Nathan Hjelm	1ae39753dc	bcol/basesmuma: check the return code of bcol_basesmuma_smcm_allgather_connection. Fixes a segmentation fault found by the bogus intercomm_create test. cmr=v1.7.4:review=manjugv This commit was SVN r30527.	2014-01-31 22:20:25 +00:00
Adrian Reber	7de34ea201	SNAPC/CRCP/SSTORE: remove compiler warnings This commit was SVN r30488.	2014-01-29 20:52:00 +00:00
Adrian Reber	fa1036f38c	SSTORE/CRCP: use ORTE_WAIT_FOR_COMPLETION with non-blocking receives During the commits to make the C/R code compile again the blocking receive calls were replaced by non-blocking which broke the code. This patch uses ORTE_WAIT_FOR_COMPLETION() to wait until the non-blocking calls have finished. This commit was SVN r30486.	2014-01-29 20:30:35 +00:00
Hadi Montakhabi	7bf4c425ff	Fix: making sure the file type is not overwritten by the last queried component This commit was SVN r30478.	2014-01-29 19:21:03 +00:00
Nathan Hjelm	afae924e29	coll/ml: fix some warnings and the spelling of indices This commit fixes one warning that should have caused coll/ml to segfault on reduce. The fix should be correct but we will continue to investigate. cmr=v1.7.5:ticket=trac:4158 This commit was SVN r30477. The following Trac tickets were found above: Ticket 4158 --> https://svn.open-mpi.org/trac/ompi/ticket/4158	2014-01-29 18:44:21 +00:00
Nathan Hjelm	700e97cf6a	btl/vader: add support for SGI's implementation of xpmem and add support for 32-bit architectures. This commit also modifies _OMPI_CHECK_HEADER to use AC_CHECK_HEADERS instead of AC_CHECK_HEADER. This allows components to check for multiple headers instead of just one. The new semantics of the header check in OMPI_CHECK_PACKAGE are to return success if at least one of the specified headers exists. The new semantics will not break current usage. cmr=v1.7.5:ticket=trac:4053 This commit was SVN r30476. The following Trac tickets were found above: Ticket 4053 --> https://svn.open-mpi.org/trac/ompi/ticket/4053	2014-01-29 18:35:47 +00:00
Jeff Squyres	3fa9d36aba	Per http://www.open-mpi.org/community/lists/devel/2014/01/13938.php , Orion Poplawski noticed that we should not be installing mpio.h. cmr=v1.7.4:reviewer=hjelmn:subject=do not install mpio.h This commit was SVN r30465.	2014-01-28 21:46:26 +00:00
George Bosilca	bde9619386	Various minor cleanups. This commit was SVN r30431.	2014-01-26 17:27:12 +00:00
George Bosilca	d265981c55	Don't always retain the proc, do it only for new procs. This enforce a strict policy in the BML, it has one and only one ref on each proc. This commit was SVN r30429.	2014-01-26 17:26:04 +00:00
Ralph Castain	b32556e6dc	Fixes trac:4143 After IM with Nathan, apply patch from ticket after verification by Paul Hargrove that it fixes the problem on non-x86 32-bit platforms Verified by Paul, RM-approved cmr=v1.7.4:reviewer=ompi-gk1.7 This commit was SVN r30411. The following Trac tickets were found above: Ticket 4143 --> https://svn.open-mpi.org/trac/ompi/ticket/4143	2014-01-24 17:56:52 +00:00
Nathan Hjelm	2435057a57	ignore the iboffload component for now. This commit was SVN r30398.	2014-01-23 16:06:21 +00:00
Rolf vandeVaart	9f3bf4747d	Provide option to have synchronous copy be asynchronous with a wait. For now, this has to be selected at runtime. Also fix up some error messages to have node name in them. This commit was SVN r30396.	2014-01-23 15:47:20 +00:00
Jeff Squyres	9fee7c2b4d	According to a report from Adam Moody, there is a compile error with ROMIO and Lustre 2.4.0. It has been solved upstream already; here's the ticket: http://trac.mpich.org/projects/mpich/ticket/1973 And here's the commit that fixed it: `a0c4278f14` OMPI does not have the other code referred to in that git commit (in ad_lustre_hints.c). Thanks to Adam Moody for reporting the issue. cmr=v1.7.4:reviewer=hjelmn:subject=Fix ROMIO compile error w/ Lustre 2.4 This commit was SVN r30393.	2014-01-23 14:15:35 +00:00
Christoph Niethammer	86776daf75	Fixed typo in opal output message. This commit was SVN r30392.	2014-01-23 08:37:40 +00:00
Mike Dubman	071838bb0a	HCOLL: call hcoll_finalize and hcoll progress unregister in case of hcoll module query failures fixed by Elena, reviewed by Val/Miked cmr=v1.7.4:reviewer=ompi-rm1.7 This commit was SVN r30390.	2014-01-23 07:29:23 +00:00
Ralph Castain	06e6a06f3e	Cleanup a couple of abstraction breaks found by Thomas Naughton This commit was SVN r30371.	2014-01-22 21:36:24 +00:00
Hadi Montakhabi	8af6b8b4e4	add support for PLFS filesystem This commit was SVN r30370.	2014-01-22 21:16:15 +00:00
Nathan Hjelm	7ba8bd81fa	coll/ml: remove debug fprintfs cmr=v1.7.5:ticket=trac:4158 This commit was SVN r30367. The following Trac tickets were found above: Ticket 4158 --> https://svn.open-mpi.org/trac/ompi/ticket/4158	2014-01-22 17:21:05 +00:00
Nathan Hjelm	82d996fb76	coll/ml: cleanup some merge related errors cmr=v1.7.5:ticket=trac:4158 This commit was SVN r30366. The following Trac tickets were found above: Ticket 4158 --> https://svn.open-mpi.org/trac/ompi/ticket/4158	2014-01-22 16:48:09 +00:00
Nathan Hjelm	ff4c9c808a	btl/ugni: fix leak in new sendi function. cmr=v1.7.5:ticket=trac:4151 This commit was SVN r30365. The following Trac tickets were found above: Ticket 4151 --> https://svn.open-mpi.org/trac/ompi/ticket/4151	2014-01-22 16:32:07 +00:00
Nathan Hjelm	66b69da394	Fix a bug in the ob1 optimizations that can cause a segfault. btl sendi functions currently can not handle the descriptor being NULL. The send inline optimization was assuming (incorrectly) that NULL was ok. cmr=v1.7.5:ticket=trac:4149 This commit was SVN r30364. The following Trac tickets were found above: Ticket 4149 --> https://svn.open-mpi.org/trac/ompi/ticket/4149	2014-01-22 16:31:58 +00:00
Nathan Hjelm	1a021b8f2d	coll/ml: add support for blocking and non-blocking allreduce, reduce, and allgather. The new collectives provide a signifigant performance increase over tuned for small and medium messages. We are initially setting the priority lower than tuned until this has had some time to soak in the trunk. Please set coll_ml_priority to 90 for MTT runs. Credit for this work goes to Manjunath Gorentla Venkata (ORNL), Pavel Shamis (ORNL), and Nathan Hjelm (LANL). Commit details (for reference): Import ORNL's collectives for MPI_Allreduce, MPI_Reduce, and MPI_Allgather. We need to take the basesmuma header into account when calculating the ptpcoll small message thresholds. Add a define to bcol.h indicating the maximum header size so we can take the header into account while not making ptpcoll dependent on information from basesmuma. This resolves an issue with allreduce where ptpcoll overwrites the header of the next buffer in the basesmuma bank. Fix reduce and make a sequential collective launcher in coll_ml_inlines.h The root calculation for reduce was wrong for any root != 0. There are four possibilities for the root: - The root is not the current process but is in the current hierarchy. In this case the root is the index of the global root as specified in the root vector. - The root is not the current process and is not in the next level of the hierarchy. In this case 0 must be the local root since this process will never communicate with the real root. - The root is not the current process but will be in next level of the hierarchy. In this case the current process must be the root. - I am the root. The root is my index. Tested with IMB which rotates the root on every call to MPI_Reduce. Consider IMB the reproducer for the issue this commit solves. Make the bcast algorithm decision an enumerated variable Resolve various asset failures when destructing coll ml requests. Two issues: - Always reset the request to be invalid before returning it to the free list. This will avoid an asset in ompi_request_t's destructor. OMPI_REQUEST_FINI does this (and also releases the fortran handle index). - Never explicitly construct or destruct the superclass of an opal object. This screws up the class function tables and will cause either an assert failure or a segmentation fault when destructing coll ml requests. Cleanup allgather. I removed the duplicate non-blocking and blocking functions and modeled the cleanup after what I found in allreduce. Also cleaned up the code somewhat. Don't bother copying from the send to the recieve buffer in bcol_basesmuma_allreduce_intra_fanin_fanout if the pointers are the same. The eliminates a warning about memcpy and aliasing and avoids an unnecessary call to memcpy. Alwasy call CHECK_AND_RELEASE on memsync collectives. There was a call to OBJ_RELEASE on the collective communicator but because CHECK_AND_RECYLCE was never called there was not matching call to OBJ_RELEASE. This caused coll ml to leak communicators. Make allreduce use the sequential collective launcher in coll_ml_inlines.h Just launch the next collective in the component progress. I am a little unsure about this patch. There appears to be some sort of race between collectives that causes buffer exhaustion in some cases (IMB Allreduce is a reproducer). Changing progress to only launch the next bcol seems to resolve the issue but might not be the best fix. Note that I see little-no performance penalty for this change. Fix allreduce when there are extra sources. There was an issue with the buffer offset calculation when there are extra sources. In the case of extra sources == 1 the offset was set to buffer_size (just past the header of the next buffer). I adjusted the buffer size to take into accoun the maximum header size (see the earlier commit that added this) and simplified the offset calculation. Make reduce/allreduce non-blocking. This is required for MPI_Comm_idup to work correctly. This has been tested with various layouts using the ibm testsuite and imb and appears to have the same performance as the old blocking version. Fix allgather for non-contiguous layouts and simplify parsing the topology. Some things in this patch: - There were several comments to the effect that level 0 of the hierarchy MUST contain all of the ranks. At least one function made this assumption but it was not true. I changed the sbgp components and the coll ml initization code to enforce this requirement. - Ensure that hierarchy level 0 has the ranks in the correct scatter gather order. This removes the need for a separate sort list and fixes the offset calculation for allgather. - There were several passes over the hierarchy to determine properties of the hierarchy. I eliminated these extra passes and the memory allocation associated with them and calculate the tree properties on the fly. The same DFS recursion also handles the re-order of level 0. All these changes have been verified with MPI_Allreduce, MPI_Reduce, and MPI_Allgather. All functions now pass all IBM/Open MPI, and IMB tests. coll/ml: correct pointer usage for MPI_BOTTOM Since contiguous datatypes are copied via memcpy (bypassing the convertor) we need to adjust for the lb of the datatype. This corrects problems found testing code that uses MPI_BOTTOM (NULL) as the send pointer. Add fallback collectives for allreduce and reduce. cmr=v1.7.5:reviewer=pasha This commit was SVN r30363.	2014-01-22 15:39:19 +00:00
Nathan Hjelm	c9c335544e	btl/ugni: fix a typo in r30353 cmr=v1.7.5:ticket=trac:4151 This commit was SVN r30354. The following SVN revision numbers were found above: r30353 --> open-mpi/ompi@aa3fea55b2 The following Trac tickets were found above: Ticket 4151 --> https://svn.open-mpi.org/trac/ompi/ticket/4151	2014-01-21 21:02:28 +00:00
Nathan Hjelm	aa3fea55b2	btl/ugni: re-add a sendi function to exploit the new optimization in ob1. Also update LANL platform files to use the latest version of ugni. cmr=v1.7.5:reviewer=manjugv This commit was SVN r30353.	2014-01-21 20:53:35 +00:00
Nathan Hjelm	2b57f4227e	ob1: optimize blocking send and receive paths Per RFC. There are two optimizations in this commit: - Allocate requests for blocking sends and receives on the stack. This bypasses the request free list and saves two atomics on the critical path. This change improves the small message ping-pong by 50-200ns on both AMD and Intel CPUs. - For small messages try to use the btl sendi function before intializing a send request. If the sendi fails or the btl does not have a sendi function silently fallback on the standard send path. cmr=v1.7.5:reviewer=brbarret This commit was SVN r30343.	2014-01-21 15:16:21 +00:00
Mike Dubman	b8550a55a7	HCOLL: many fixes Adds coll_hcoll_np mca parameter similar to that of fca component (defaults to 32). Those who use hcoll be aware that from now on the communicators less than 32 procs will run w/o hcoll by default. - Resolves fallback issue in case libhcoll runs out of allowed contexts. The solution is moving hcoll_context_create from comm_enable to comm_query. Shortly, comm_enable should never return OMPI_ERROR in the coll component with highest priority (hcoll). Otherwise the ompi coll_base_select will unselect the coll funtion pointers and module references leaving the communicator w/o coll pointer. This will cause the fail. Same behavior can be reproduced even with tuned if one would hardcore some "return OMPI_ERROR" into it's module_enable funtion. - Additionally, removed all the dead code under #if 0; removed unused variables (path for library, active_modules list) and classes (module list wrapper) Fixed by Val, Reviewed by Devendar/Josh/Miked cmr=v1.7.4:reviewer=ompi-rm1.7 This commit was SVN r30341.	2014-01-21 12:19:47 +00:00
Ralph Castain	2cf4862b49	Cleanup warnings for use of void* - requires intermediate cast to uintptr_t. Thanks to Paul Hargrove for reporting it cmr=v1.7.4:reviewer=jsquyres This commit was SVN r30333.	2014-01-20 15:44:45 +00:00
Edgar Gabriel	be5d5834c5	fix the problem identified by a user on the mailing list with MPI_MODE_EXCL cmr=v1.7.4:reviewer=vvenkatesan:subject=fix a problem when opening a file with MODE_EXCL This commit was SVN r30324.	2014-01-18 16:06:27 +00:00
Nathan Hjelm	c88626510c	Fix a merge issues with new ROMIO and fix obvious ROMIO bug. cmr=v1.7.4:reviewer=jsquyres This commit was SVN r30319.	2014-01-18 00:29:16 +00:00
Hadi Montakhabi	8c14411289	f_cc_size is contiguous chunk size, not the stripe width. There is no stripe_width in the file handle structure. This commit was SVN r30314.	2014-01-17 18:35:55 +00:00
Nathan Hjelm	f2a73fcdbd	udreg: free huge page allocations correctly This commit fixes an error path that occurs when huge page allocations are enabled. In this case we allocate a huge page and try to register it but fail. We then were calling free on the opal object. Fix this by calling the proper free function. cmr=v1.7.4:reviewer=rhc This commit was SVN r30289.	2014-01-14 16:26:06 +00:00
Nathan Hjelm	f9d2032705	vader: ensure fast box data is aligned on 4-byte boundaries This commit fixes a bus error on Solaris/Sparc. Closes trac:4111 cmr=v1.7.5:ticket=trac:4053 This commit was SVN r30288. The following Trac tickets were found above: Ticket 4053 --> https://svn.open-mpi.org/trac/ompi/ticket/4053 Ticket 4111 --> https://svn.open-mpi.org/trac/ompi/ticket/4111	2014-01-14 16:04:52 +00:00
Rolf vandeVaart	e75afb2b82	Fix bug in distance computation code when deciding which devices to use on a NUMA node. Also add a verbose flag so one can see what devices are selected as well as another flag to override locality information and use all devices on the node. This commit was SVN r30287.	2014-01-14 15:41:56 +00:00
Nathan Hjelm	da1316ca6e	vader: don't OBJ_RELEASE endpoint rcaches. cmr=v1.7.4:reviewer=rhc This commit was SVN r30284.	2014-01-13 23:44:34 +00:00
Jeff Squyres	20d6391734	Patch submitted by Paul Hargrove to fix NetBSD compile with -laio. NetBSD puts the AIO functions in -lrt, vs. the usual libc. So we need the fbtl/posix configure.m4 to test for -lrt properly. Reviewed by Jeff Squyres. cmr=v1.7.4:reviewer=ompi-rm1.7:subject=Fix NetBSD use of -laio This commit was SVN r30274.	2014-01-13 18:49:39 +00:00
Yossi Etigin	7564e2c13f	Fix a recursion in mxm send flow which happens when mpi starts a new send from the context of send completion callback. cmr=v1.7.5:reviewer=jsquyres This commit was SVN r30265.	2014-01-12 17:47:03 +00:00
Yossi Etigin	9504969f7d	fix communicator double-free from pt2pt component, caused by r29938. cmr=v1.7.5:reviewer=brbarret This commit was SVN r30264. The following SVN revision numbers were found above: r29938 --> open-mpi/ompi@ecfb122c97	2014-01-12 17:38:14 +00:00
Ralph Castain	286ff6d552	For large scale systems, we would like to avoid doing a full modex during MPI_Init so that launch will scale a little better. At the moment, our options are somewhat limited as only a few BTLs don't immediately call modex_recv on all procs during startup. However, for those situations where someone can take advantage of it, add the ability to do a "modex on demand" retrieval of data from remote procs when we launch via mpirun. NOTE: launch performance will be absolutely awful if you do this with BTLs that aren't configured to modex_recv on first message! Even with "modex on demand", we still have to do a barrier in place of the modex - we simply don't move any data around, which does reduce the time impact. The barrier is required to ensure that the other proc has in fact registered all its BTL info and therefore is prepared to hand over a complete data package. Otherwise, you may not get the info you need. In addition, the shared memory BTL can fail to properly rendezvous as it expects the barrier to be in place. This behavior will only take effect under the following conditions: 1. launched via mpirun 2. #procs is greater than ompi_hostname_cutoff, which defaults to UINT32_MAX 3. mca param rte_orte_direct_modex is set to 1. At the moment, we are having problems getting this param to register properly, so only the first two conditions are in effect. Still, the bottom line is you have to want this behavior to get it. The planned next evolution of this will be to make the direct modex be non-blocking - this will require two fixes: 1. if the remote proc doesn't have the required info, then let it delay its response until it does. This means we need a way for the MPI layer to tell the RTE "I am done entering modex data". 2. adjust the SM rendezvous logic to loop until the required file has been created Creating a placeholder to bring this over to 1.7.5 when ready. cmr=v1.7.5:reviewer=hjelmn:subject=Enable direct modex at scale This commit was SVN r30259.	2014-01-11 17:36:06 +00:00

1 2 3 4 5 ...

4610 Коммитов