openmpi

Автор	SHA1	Сообщение	Дата
Jeff Squyres	8ace07efed	This commit brings in two major things: 1. Galen's fine-grain control of queue pair resources in the openib BTL. 1. Pasha's new implementation of asychronous HCA event handling. Pasha's new implementation doesn't take much explanation, but the new "multifrag" stuff does. Note that "svn merge" was not used to bring this new code from the /tmp/ib_multifrag branch -- something Bad happened in the periodic trunk pulls on that branch making an actual merge back to the trunk effectively impossible (i.e., lots and lots of arbitrary conflicts and artifical changes). :-( == Fine-grain control of queue pair resources == Galen's fine-grain control of queue pair resources to the OpenIB BTL (thanks to Gleb for fixing broken code and providing additional functionality, Pasha for finding broken code, and Jeff for doing all the svn work and regression testing). Prior to this commit, the OpenIB BTL created two queue pairs: one for eager size fragments and one for max send size fragments. When the use of the shared receive queue (SRQ) was specified (via "-mca btl_openib_use_srq 1"), these QPs would use a shared receive queue for receive buffers instead of the default per-peer (PP) receive queues and buffers. One consequence of this design is that receive buffer utilization (the size of the data received as a percentage of the receive buffer used for the data) was quite poor for a number of applications. The new design allows multiple QPs to be specified at runtime. Each QP can be setup to use PP or SRQ receive buffers as well as giving fine-grained control over receive buffer size, number of receive buffers to post, when to replenish the receive queue (low water mark) and for SRQ QPs, the number of outstanding sends can also be specified. The following is an example of the syntax to describe QPs to the OpenIB BTL using the new MCA parameter btl_openib_receive_queues: {{{ -mca btl_openib_receive_queues \ "P,128,16,4;S,1024,256,128,32;S,4096,256,128,32;S,65536,256,128,32" }}} Each QP description is delimited by ";" (semicolon) with individual fields of the QP description delimited by "," (comma). The above example therefore describes 4 QPs. The first QP is: P,128,16,4 Meaning: per-peer receive buffer QPs are indicated by a starting field of "P"; the first QP (shown above) is therefore a per-peer based QP. The second field indicates the size of the receive buffer in bytes (128 bytes). The third field indicates the number of receive buffers to allocate to the QP (16). The fourth field indicates the low watermark for receive buffers at which time the BTL will repost receive buffers to the QP (4). The second QP is: S,1024,256,128,32 Shared receive queue based QPs are indicated by a starting field of "S"; the second QP (shown above) is therefore a shared receive queue based QP. The second, third and fourth fields are the same as in the per-peer based QP. The fifth field is the number of outstanding sends that are allowed at a given time on the QP (32). This provides a "good enough" mechanism of flow control for some regular communication patterns. QPs MUST be specified in ascending receive buffer size order. This requirement may be removed prior to 1.3 release. This commit was SVN r15474.	2007-07-18 01:15:59 +00:00
Galen Shipman	3401bd2b07	Add optional ordering to the BTL interface. This is required to tighten up the BTL semantics. Ordering is not guaranteed, but, if the BTL returns a order tag in a descriptor (other than MCA_BTL_NO_ORDER) then we may request another descriptor that will obey ordering w.r.t. to the other descriptor. This will allow sane behavior for RDMA networks, where local completion of an RDMA operation on the active side does not imply remote completion on the passive side. If we send a FIN message after local completion and the FIN is not ordered w.r.t. the RDMA operation then badness may occur as the passive side may now try to deregister the memory and the RDMA operation may still be pending on the passive side. Note that this has no impact on networks that don't suffer from this limitation as the ORDER tag can simply always be specified as MCA_BTL_NO_ORDER. This commit was SVN r14768.	2007-05-24 19:51:26 +00:00
Gleb Natapov	90fb58de4f	When frags are allocated from mpool by free_list the frag structure is also allocated from mpool memory (which is registered memory for RDMA transports) This is not a problem for a small jobs, but for a big number of ranks an amount of waisted memory is big. This commit was SVN r13921.	2007-03-05 14:17:50 +00:00
Gleb Natapov	2b6cbd6299	Separate frag lists for RDMA descriptors to two, one for src descriptors and another for dst descriptors. This provide partial solution to OB1 protocol deadlock problem. We can limit number of RDMA descriptors (by setting btl_openib_free_list_max to something different from -1) and if we will be lucky to hit this limit before we fail to register more memory the protocol will not deadlock. When we had only one list for src/dst descriptors we deadlocked when we reached max limit for the list. This commit was SVN r13844.	2007-02-28 13:43:38 +00:00
Gleb Natapov	624f139bd8	This commit fixes trac:729. Initialize pointer to registration to NULL. Otherwise it may contain garbage and we will try to unregister it later in btl_free(). This commit was SVN r13054. The following Trac tickets were found above: Ticket 729 --> https://svn.open-mpi.org/trac/ompi/ticket/729	2007-01-09 10:29:20 +00:00
Brian Barrett	48ec0b2071	Revert out r12974, 12976, and 12991 as George has provided a less intrusive fix for now... This commit was SVN r12997. The following SVN revision numbers were found above: r12974 --> open-mpi/ompi@27cea44a9c	2007-01-04 22:07:37 +00:00
Brian Barrett	27cea44a9c	Fix a number of issues with the ompi_ptr_t: * Make sure that the pval always writes to the correct portion of the lval. This only matters on 32 bit big endian machines. * On 32 bit machines when assigning to pval, the other 4 bytes of lval weren't being written, which could lead to bogus data We use macros so that there aren't casts all over the code and the pval assignment can occur to the correct 4 bytes. Refs trac:587 This commit was SVN r12974. The following Trac tickets were found above: Ticket 587 --> https://svn.open-mpi.org/trac/ompi/ticket/587	2007-01-03 19:47:48 +00:00
Gleb Natapov	190e7a27cd	Merge with gleb-mpool branch. All RDMA components use same mpool now (rdma). udapl/openib/vapi/gm mpools a deprecated. rdma mpool has parameter that allows to limit its size mpool_rdma_rcache_size_limit (default is 0 - unlimited). This commit was SVN r12878.	2006-12-17 12:26:41 +00:00
Gleb Natapov	7999c08107	consolidate credit management and CQ polling code. This commit was SVN r11622.	2006-09-12 09:17:59 +00:00
Gleb Natapov	d0caffa0aa	Consolidate receive buffers prepost code for HP/LP QPs. This commit was SVN r11552.	2006-09-07 13:05:41 +00:00
Gleb Natapov	72575d81d2	Create separate pool for control messages. It is unlimited, but the maximum number of element that are allocated from it is limited by number of connections. This commit was SVN r11028.	2006-07-27 14:09:30 +00:00
Gleb Natapov	52208d7bf9	Whe don't need to register zero sized frags. This commit was SVN r10519.	2006-06-27 08:50:12 +00:00
Gleb Natapov	79bcfb096f	Add type to frag. Sometimes we need to know that a frag is from short rdma area. I used hack for this that doesn't work for mvapi, so changing it to something more sane. This commit was SVN r9477.	2006-03-30 15:26:21 +00:00
Gleb Natapov	a5a78b10cc	Implementation of short message RDMA. Endpoint registers circular buffer and sends its address and rkey to the peer. Peer uses this buffer to eagerly RDMA small message into it. Endpoint polls the buffer for message arrival before checking HP/LP QPs. Set btl_openib_use_eager_rdma to 1 to enable it. This commit was SVN r9425.	2006-03-26 08:30:50 +00:00
Brian Barrett	566a050c23	Next step in the project split, mainly source code re-arranging - move files out of toplevel include/ and etc/, moving it into the sub-projects - rather than including config headers with <project>/include, have them as <project> - require all headers to be included with a project prefix, with the exception of the config headers ({opal,orte,ompi}_config.h mpi.h, and mpif.h) This commit was SVN r8985.	2006-02-12 01:33:29 +00:00
Tim Woodall	a584c60dbe	re-worked flow control logic to take into account the return of credits from the peer prior to local completion, so that we don't overrun the number of send wqes available. This commit was SVN r8683.	2006-01-12 23:42:44 +00:00
Galen Shipman	635e7a682b	fix for 32bit compile warnings. This commit was SVN r8190.	2005-11-18 17:08:51 +00:00
Galen Shipman	dde38d4119	reset sg_entry->addr to point at header when sending control messages. cast to uint64_t (the correct datatype per verbs.h) instead of uintptr_t. This commit was SVN r8175.	2005-11-17 05:45:33 +00:00
Tim Woodall	4a06e8463c	port of flow control from mvapi This commit was SVN r8102.	2005-11-10 20:15:02 +00:00
Jeff Squyres	42ec26e640	Update the copyright notices for IU and UTK. This commit was SVN r7999.	2005-11-05 19:57:48 +00:00
Galen Shipman	67d38b7896	Add multi-nic support to openib Fix connection establishment race in openib Other misc This commit was SVN r7570.	2005-09-30 22:58:09 +00:00
Galen Shipman	946402b980	More openib cleanup.. still note ready for public consumption ;-) This commit was SVN r6565.	2005-07-20 15:17:18 +00:00
Galen Shipman	2f67ab82bb	Working version of openib btl ;-) Fixed receive descriptor counts that limited mvapi and openib to 2 procs. Begin porting error messages to use the BTL_ERROR macro. This commit was SVN r6554.	2005-07-19 21:04:22 +00:00
Galen Shipman	d7bdc46ac9	compile error and warining fixes for openib.. This commit was SVN r6449.	2005-07-12 21:49:30 +00:00
Galen Shipman	454fdff824	Initial commit of changes to the mvapi btl to the openib btl. Still need to work on the configure.stub to correctly locate the ib libraries. This commit was SVN r6435.	2005-07-12 13:38:54 +00:00
Jeff Squyres	4ab17f019b	Rename src -> ompi This commit was SVN r6269.	2005-07-02 13:43:57 +00:00

26 Коммитов