openmpi

Автор	SHA1	Сообщение	Дата
Ralph Castain	9613b3176c	Effectively revert the orte_output system and return to direct use of opal_output at all levels. Retain the orte_show_help subsystem to allow aggregation of show_help messages at the HNP. After much work by Jeff and myself, and quite a lot of discussion, it has become clear that we simply cannot resolve the infinite loops caused by RML-involved subsystems calling orte_output. The original rationale for the change to orte_output has also been reduced by shifting the output of XML-formatted vs human readable messages to an alternative approach. I have globally replaced the orte_output/ORTE_OUTPUT calls in the code base, as well as the corresponding .h file name. I have test compiled and run this on the various environments within my reach, so hopefully this will prove minimally disruptive. This commit was SVN r18619.	2008-06-09 14:53:58 +00:00
Gleb Natapov	5fabade090	Use payload_buffer_alignment value for payload alignment. This commit was SVN r18493.	2008-05-26 08:29:02 +00:00
Jeff Squyres	e7ecd56bd2	This commit represents a bunch of work on a Mercurial side branch. As such, the commit message back to the master SVN repository is fairly long. = ORTE Job-Level Output Messages = Add two new interfaces that should be used for all new code throughout the ORTE and OMPI layers (we already make the search-and-replace on the existing ORTE / OMPI layers): * orte_output(): (and corresponding friends ORTE_OUTPUT, orte_output_verbose, etc.) This function sends the output directly to the HNP for processing as part of a job-specific output channel. It supports all the same outputs as opal_output() (syslog, file, stdout, stderr), but for stdout/stderr, the output is sent to the HNP for processing and output. More on this below. * orte_show_help(): This function is a drop-in-replacement for opal_show_help(), with two differences in functionality: 1. the rendered text help message output is sent to the HNP for display (rather than outputting directly into the process' stderr stream) 1. the HNP detects duplicate help messages and does not display them (so that you don't see the same error message N times, once from each of your N MPI processes); instead, it counts "new" instances of the help message and displays a message every ~5 seconds when there are new ones ("I got X new copies of the help message...") opal_show_help and opal_output still exist, but they only output in the current process. The intent for the new orte_* functions is that they can apply job-level intelligence to the output. As such, we recommend that all new ORTE and OMPI code use the new orte_* functions, not thei opal_* functions. === New code === For ORTE and OMPI programmers, here's what you need to do differently in new code: * Do not include opal/util/show_help.h or opal/util/output.h. Instead, include orte/util/output.h (this one header file has declarations for both the orte_output() series of functions and orte_show_help()). * Effectively s/opal_output/orte_output/gi throughout your code. Note that orte_output_open() takes a slightly different argument list (as a way to pass data to the filtering stream -- see below), so you if explicitly call opal_output_open(), you'll need to slightly adapt to the new signature of orte_output_open(). * Literally s/opal_show_help/orte_show_help/. The function signature is identical. === Notes === * orte_output'ing to stream 0 will do similar to what opal_output'ing did, so leaving a hard-coded "0" as the first argument is safe. * For systems that do not use ORTE's RML or the HNP, the effect of orte_output_* and orte_show_help will be identical to their opal counterparts (the additional information passed to orte_output_open() will be lost!). Indeed, the orte_* functions simply become trivial wrappers to their opal_* counterparts. Note that we have not tested this; the code is simple but it is quite possible that we mucked something up. = Filter Framework = Messages sent view the new orte_* functions described above and messages output via the IOF on the HNP will now optionally be passed through a new "filter" framework before being output to stdout/stderr. The "filter" OPAL MCA framework is intended to allow preprocessing to messages before they are sent to their final destinations. The first component that was written in the filter framework was to create an XML stream, segregating all the messages into different XML tags, etc. This will allow 3rd party tools to read the stdout/stderr from the HNP and be able to know exactly what each text message is (e.g., a help message, another OMPI infrastructure message, stdout from the user process, stderr from the user process, etc.). Filtering is not active by default. Filter components must be specifically requested, such as: {{{ $ mpirun --mca filter xml ... }}} There can only be one filter component active. = New MCA Parameters = The new functionality described above introduces two new MCA parameters: * '''orte_base_help_aggregate''': Defaults to 1 (true), meaning that help messages will be aggregated, as described above. If set to 0, all help messages will be displayed, even if they are duplicates (i.e., the original behavior). * '''orte_base_show_output_recursions''': An MCA parameter to help debug one of the known issues, described below. It is likely that this MCA parameter will disappear before v1.3 final. = Known Issues = * The XML filter component is not complete. The current output from this component is preliminary and not real XML. A bit more work needs to be done to configure.m4 search for an appropriate XML library/link it in/use it at run time. * There are possible recursion loops in the orte_output() and orte_show_help() functions -- e.g., if RML send calls orte_output() or orte_show_help(). We have some ideas how to fix these, but figured that it was ok to commit before feature freeze with known issues. The code currently contains sub-optimal workarounds so that this will not be a problem, but it would be good to actually solve the problem rather than have hackish workarounds before v1.3 final. This commit was SVN r18434.	2008-05-13 20:00:55 +00:00
Tim Prins	889f6c79fe	Properly initialize the freelist in the constructor. This commit was SVN r17175.	2008-01-22 18:17:06 +00:00
Rich Graham	e4646a4dd5	going through the ompi_free_list_init_ex, fl_payload_buffer_size and fl_payload_buffer_alignment were not being set. This commit was SVN r16641.	2007-11-02 17:51:32 +00:00
Rich Graham	27a748e7eb	change all instances of ompi_free_list_init to ompi_free_list_init_new. Header and payload data are specified separately at this stage. This commit was SVN r16633.	2007-11-01 23:38:50 +00:00
Rich Graham	aa82acd34c	continuing the incremental changes. fl_elem_class renamed fl_frag_class, and ompi_free_list_init_new() and ompi_free_list_init_ex_new() were added. Next step will be to start converting from ompi_free_list_init to() ompi_free_list_init_new(), and then remove ompi_free_list_init(), and rename ompi_free_list_init_new() back to ompi_free_list_init(). The merge of the branch with the trunk was so substantial, it is far easeir to re-implement the changes in the trunk, rather than trying to fix the bugs the merge brought in ... This commit was SVN r16630.	2007-11-01 17:25:12 +00:00
Rich Graham	52fb318950	starting to put in the changes for ompi_free_list_t. fl_elem_size is renamed to fl_frag_size, fl_alignment is renamed to fl_frag_alignment, and fl_payload_buffer_size and fl_payload_buffer_alignment are added. This commit was SVN r16629.	2007-11-01 16:47:44 +00:00
George Bosilca	e724ca0a1f	Remodel the ompi_free_list a little. The free_list_memory is in fact a free_list_item so instead of having a struct, use typedef to make them equivalent. Modify the parallel debuggers support in order to allow them access to the internal types even when we have an optimized build. This commit was SVN r16567.	2007-10-25 16:47:54 +00:00
Shiqing Fan	a0660f4deb	- Just some type casts. This commit was SVN r16100.	2007-09-12 15:29:58 +00:00
Jeff Squyres	8ace07efed	This commit brings in two major things: 1. Galen's fine-grain control of queue pair resources in the openib BTL. 1. Pasha's new implementation of asychronous HCA event handling. Pasha's new implementation doesn't take much explanation, but the new "multifrag" stuff does. Note that "svn merge" was not used to bring this new code from the /tmp/ib_multifrag branch -- something Bad happened in the periodic trunk pulls on that branch making an actual merge back to the trunk effectively impossible (i.e., lots and lots of arbitrary conflicts and artifical changes). :-( == Fine-grain control of queue pair resources == Galen's fine-grain control of queue pair resources to the OpenIB BTL (thanks to Gleb for fixing broken code and providing additional functionality, Pasha for finding broken code, and Jeff for doing all the svn work and regression testing). Prior to this commit, the OpenIB BTL created two queue pairs: one for eager size fragments and one for max send size fragments. When the use of the shared receive queue (SRQ) was specified (via "-mca btl_openib_use_srq 1"), these QPs would use a shared receive queue for receive buffers instead of the default per-peer (PP) receive queues and buffers. One consequence of this design is that receive buffer utilization (the size of the data received as a percentage of the receive buffer used for the data) was quite poor for a number of applications. The new design allows multiple QPs to be specified at runtime. Each QP can be setup to use PP or SRQ receive buffers as well as giving fine-grained control over receive buffer size, number of receive buffers to post, when to replenish the receive queue (low water mark) and for SRQ QPs, the number of outstanding sends can also be specified. The following is an example of the syntax to describe QPs to the OpenIB BTL using the new MCA parameter btl_openib_receive_queues: {{{ -mca btl_openib_receive_queues \ "P,128,16,4;S,1024,256,128,32;S,4096,256,128,32;S,65536,256,128,32" }}} Each QP description is delimited by ";" (semicolon) with individual fields of the QP description delimited by "," (comma). The above example therefore describes 4 QPs. The first QP is: P,128,16,4 Meaning: per-peer receive buffer QPs are indicated by a starting field of "P"; the first QP (shown above) is therefore a per-peer based QP. The second field indicates the size of the receive buffer in bytes (128 bytes). The third field indicates the number of receive buffers to allocate to the QP (16). The fourth field indicates the low watermark for receive buffers at which time the BTL will repost receive buffers to the QP (4). The second QP is: S,1024,256,128,32 Shared receive queue based QPs are indicated by a starting field of "S"; the second QP (shown above) is therefore a shared receive queue based QP. The second, third and fourth fields are the same as in the per-peer based QP. The fifth field is the number of outstanding sends that are allowed at a given time on the QP (32). This provides a "good enough" mechanism of flow control for some regular communication patterns. QPs MUST be specified in ascending receive buffer size order. This requirement may be removed prior to 1.3 release. This commit was SVN r15474.	2007-07-18 01:15:59 +00:00
Gleb Natapov	90fb58de4f	When frags are allocated from mpool by free_list the frag structure is also allocated from mpool memory (which is registered memory for RDMA transports) This is not a problem for a small jobs, but for a big number of ranks an amount of waisted memory is big. This commit was SVN r13921.	2007-03-05 14:17:50 +00:00
Pavel Shamis	edeab0e912	Adding Mellanox Technologies copyright to files touched by Mellanox. This commit was SVN r13669.	2007-02-15 18:03:20 +00:00
Gleb Natapov	190e7a27cd	Merge with gleb-mpool branch. All RDMA components use same mpool now (rdma). udapl/openib/vapi/gm mpools a deprecated. rdma mpool has parameter that allows to limit its size mpool_rdma_rcache_size_limit (default is 0 - unlimited). This commit was SVN r12878.	2006-12-17 12:26:41 +00:00
George Bosilca	b51b87a4aa	The correct way to compute the difference between the actual size and the expected size, based on the comment few lines before. This commit was SVN r12235.	2006-10-20 19:33:55 +00:00
George Bosilca	e81d38f322	Remove a function that was just a proof of concept. The same approach is not used by the TotalView support. This commit was SVN r12214.	2006-10-20 03:34:16 +00:00
George Bosilca	33f300f636	I don't know what it was supposed to do but I'm quite sure it didn't do it correctly. Just follow inc_num and you will understand. Now _resize will grow the list to match the required number of elements as described in the comment in the .h file. This commit was SVN r12074.	2006-10-10 14:47:51 +00:00
Pavel Shamis	e400da01bc	Fix of typo error introduced in revision 12053. Reviewed by: Jeff This commit was SVN r12071.	2006-10-10 11:58:29 +00:00
Brian Barrett	51b2a0fd3f	A couple of changes to improve shared memory behavior when resources get constrained: * Make sure we always have a number of eager fragments available that scales with the number of processes communicating with a given proc over shared memory * Use FREE_LIST_GET instead of FREE_LIST_WAIT to return an error to the PML when resource exhaustion occurs * Don't dereference the frag during alloc unless we're sure it's not NULL Reviewed by: Galen Refs trac:413 This commit was SVN r12053. The following Trac tickets were found above: Ticket 413 --> https://svn.open-mpi.org/trac/ompi/ticket/413	2006-10-06 21:13:49 +00:00
George Bosilca	3f0a7cad9e	The last patch for Windows support. Mostly casting and conversion to C++ friendly headers. This commit was SVN r11400.	2006-08-24 16:38:08 +00:00
Gleb Natapov	383694c68d	Add support to get alignemnt buffers from free_list_t. Convert openib BTL to new interface. This commit was SVN r10899.	2006-07-20 14:39:05 +00:00
George Bosilca	9eb023a5c2	OK my last commit was ... kind of wrong. It only worked if the element_size was smaller than the CACHE_LINE_SIZE. Here is the version that works. In fact this works on 2 steps. First we set the element size to something multiple of the desired alignment. Then when we allocate memory, we compute the total size, and we will align each of the elements (we allocate multiple of them every time) to the CACHE_LINE_SIZE. This commit was SVN r10479.	2006-06-22 14:47:07 +00:00
George Bosilca	c71f6c9765	All elements will be aligned to the CACHE_LINE_SIZE define (currently 128 bytes). The simplest way to make sure they are aligned is to update the size of the basic element to a multiple of the desired alignment. It will use a little bit more memory, but the improvements on the SM BTL seems quite interesting. This commit was SVN r10478.	2006-06-22 14:07:14 +00:00
George Bosilca	aca71521db	Complete the move of the mpool registration from opal_list_item_t to the ompi_free_list_item_t. This commit was SVN r10354.	2006-06-14 17:43:50 +00:00
Galen Shipman	18dda70fd0	make ompi_free_list_item_t a class.. This will go to the 1.1 branch but will probably require a few changes as ompi_free_list_t is different in the branch.. This commit was SVN r10306.	2006-06-12 16:44:00 +00:00
George Bosilca	e68382a66d	Add a new debug function. It will parse all the items alocated by this free list. It use the size attached to the free list, and the internal memory segments to find out all the items allocated by this free list. This commit was SVN r9669.	2006-04-20 19:53:45 +00:00
George Bosilca	b92e78761a	I did it under pressure !!! The free lst using atomic operations. I didn't want to completely change the behavior, so we still use a mutex for the extreme cases (like no more available items and we cannot allocate more). I test it for a while on non multi-threading environment, but not enough on a multi-threaded build. This commit was SVN r9623.	2006-04-12 23:27:38 +00:00
George Bosilca	0c58d0f519	Mixing declarations and code is not allowed by the ISO C90. This commit was SVN r9494.	2006-03-31 03:21:28 +00:00
George Bosilca	994959345a	The if outside the loop not inside as we test for a "constant" thing. This commit was SVN r9488.	2006-03-31 00:38:11 +00:00
Tim Woodall	bd870519fd	- modified convertor copy_and_prepare routines to accept an addition flag, new flags to be included when convertor is initialized - modified pml/btl module defs and added stub functions for diagnostic output routines to dump state of queues / endpoints - updates to data reliability pml This commit was SVN r9329.	2006-03-17 18:46:48 +00:00
Tim Woodall	bab5b2a63e	check for resource leak This commit was SVN r9310.	2006-03-16 17:27:54 +00:00
Brian Barrett	566a050c23	Next step in the project split, mainly source code re-arranging - move files out of toplevel include/ and etc/, moving it into the sub-projects - rather than including config headers with <project>/include, have them as <project> - require all headers to be included with a project prefix, with the exception of the config headers ({opal,orte,ompi}_config.h mpi.h, and mpif.h) This commit was SVN r8985.	2006-02-12 01:33:29 +00:00
George Bosilca	d5d16c2162	A slighy faster version. The if outside the for not inside. This commit was SVN r8761.	2006-01-19 23:57:03 +00:00
Jeff Squyres	42ec26e640	Update the copyright notices for IU and UTK. This commit was SVN r7999.	2005-11-05 19:57:48 +00:00
Tim Woodall	3bd5b81dfa	Submitted: Gleb Natapov This commit was SVN r7899.	2005-10-27 17:48:40 +00:00
Galen Shipman	d932cfd342	merge of rcache work into the trunk.. lotsa fun ;-).. I regression tested before the merge, I will regression test tonight and correct issues that might have crept in. This commit was SVN r7329.	2005-09-12 22:28:23 +00:00
Brian Barrett	492aeecd11	* track the actual allocations within the ompi and opal free_lists, so that we can free all allocated memory when the free_list is destructed. Fixes a whole bunch of memory leaks... This commit was SVN r7173.	2005-09-03 19:46:44 +00:00
Brian Barrett	a9863e5ff5	* properly construct / destruct the condition variable in the free list structures This commit was SVN r7153.	2005-09-02 16:00:42 +00:00
Galen Shipman	ba82bc11bc	bug fixes and configure check for topspin directory structure.. This commit was SVN r6767.	2005-08-08 19:10:36 +00:00
Brian Barrett	39dbeeedfb	* rename locking code from ompi to opal This commit was SVN r6327.	2005-07-03 22:45:48 +00:00
Brian Barrett	761402f95f	* rename ompi_list to opal_list This commit was SVN r6322.	2005-07-03 16:22:16 +00:00
Brian Barrett	499e4de1e7	* rename ompi_object and ompi_class to opal_object and opal_class This commit was SVN r6321.	2005-07-03 16:06:07 +00:00
Jeff Squyres	959a08bf42	Compromise: - move mpool and allocator frameworks back to ompi (from opal) - specialize the ompi_free_list class to use an mpool instance - un-specialize opal_free_list to not use mpool; just use malloc/free This commit was SVN r6292.	2005-07-02 15:35:34 +00:00

43 Коммитов