openmpi

Автор	SHA1	Сообщение	Дата
Sven Stork	6111ca1152	- Let's try to detect the default nodefile directory because it can different for different sites. If we cannot detect the default then we fall back to the hard coded path. This commit was SVN r14121.	2007-03-22 15:26:16 +00:00
Galen Shipman	e654604a25	remove invalid comment This commit was SVN r14118.	2007-03-22 03:51:36 +00:00
Josh Hursey	3492fdeae3	Fix a couple of compiler warnings (errors?) caught by ICC testing at Cisco. This commit was SVN r14080.	2007-03-20 14:12:13 +00:00
Rainer Keller	1322f9f346	- Further attributes mainly for opal/* functions, marking __opal_attribute_nonnull__, __opal_attribute_warn_unused_result__, __opal_attribute_malloc__, __opal_attribute_sentinel__ and __opal_attribute_format__ This commit was SVN r14078.	2007-03-20 13:01:32 +00:00
Pak Lui	803655b555	* incorporated some of Jeff's comment regarding this fix. This commit was SVN r14070.	2007-03-19 21:59:48 +00:00
Pak Lui	da4d41e0e7	* fixed the missing fclose and eliminate the call to get_slot_count since it is not needed This commit was SVN r14066.	2007-03-19 17:47:30 +00:00
Rich Graham	d2e799f6b5	add some stub functions for the cnos environment. This commit was SVN r14065.	2007-03-19 17:35:46 +00:00
Josh Hursey	101a2abd09	- Be more careful with parens - Run the destructor before shutting things down. This commit was SVN r14064.	2007-03-19 17:33:20 +00:00
Brian Barrett	ea08a555f9	Fixed a compile error on OS X 10.3 introduced with 1.1.5 / 1.2. Thanks to Marius Schamschula for reporting the issue. This commit was SVN r14063.	2007-03-19 17:25:54 +00:00
Josh Hursey	d03073e87d	Make sure to protect the finalize call so tools like ompi_info do not segv. This commit was SVN r14054.	2007-03-17 19:47:54 +00:00
Josh Hursey	dadca7da88	Merging in the jjhursey-ft-cr-stable branch (r13912 : HEAD). This merge adds Checkpoint/Restart support to Open MPI. The initial frameworks and components support a LAM/MPI-like implementation. This commit follows the risk assessment presented to the Open MPI core development group on Feb. 22, 2007. This commit closes trac:158 More details to follow. This commit was SVN r14051. The following SVN revisions from the original message are invalid or inconsistent and therefore were not cross-referenced: r13912 The following Trac tickets were found above: Ticket 158 --> https://svn.open-mpi.org/trac/ompi/ticket/158	2007-03-16 23:11:45 +00:00
Jeff Squyres	c000ee5328	Fixes trac:921 * Do not empty the list of in-flight frags during _close(); the OOB callback will still occur (_send_cb()) and try to remove the frag from the list, which will then result in an assert failure (debug builds). * Add one more fix for a possible problem -- add an extra RETAIN / RELEASE pair on the endpoint to ensure that it is not actually freed before all in-flight frags have drained. This commit was SVN r13953. The following Trac tickets were found above: Ticket 921 --> https://svn.open-mpi.org/trac/ompi/ticket/921	2007-03-07 20:12:22 +00:00
Jeff Squyres	7b72ded10c	Patch from Gotz Waschk to recognize zsh. This commit was SVN r13907.	2007-03-03 01:42:03 +00:00
Li-Ta Lo	a0e5b6a27c	minor clean up and treespawn support This commit was SVN r13876.	2007-03-01 22:32:37 +00:00
Josh Hursey	0404444dbe	* Added 2 new MCA parameters - mca_base_param_file_prefix (Default: NULL) This is the fullname of the "-am" mpirun option. Used to specify a ':' separated list of AMCA parameter set files. - mca_base_param_file_path (Default: $SYSCONFDIR/amca-param-sets/:$CWD) The path to search for AMCA files with relative paths. A warning will be printed if the AMCA file cannot be found. * Added a new function "mca_base_param_recache_files" the re-reads the file configurations. This is used internally to help bootstrap the MCA system. * Added a new orterun/mpirun command line option '-am' that aliases for the mca_base_param_file_prefix MCA parameter * Exposed the opal_path_access function as it is generally useful in other places in the code. * New function "opal_cmd_line_make_opt_mca" which will allow you to append a new command line option with MCA parameter identifiers to set at the same time. Previously this could only be done at command line declaration time. * Added a new directory under the $pkgdatadir named "amca-param-sets" where all the 'shipped with' Open MPI AMCA parameter sets are placed. This is the first place to search for AMCA sets with relative paths. * An example.conf AMCA parameter set file is located in contrib/amca-param-sets/. * Jeff Squyres contributed an OpenIB AMCA set for benchmarking. Note: You will need to autogen with this commit as it adds a configure param. Sorry :( This commit was SVN r13867.	2007-03-01 13:39:20 +00:00
Rainer Keller	0889ebd59f	- Eliminate warnings, that PGI-6.2.5 issues with -Minform=inform This commit was SVN r13840.	2007-02-28 08:36:34 +00:00
George Bosilca	4bab882d17	These 2 ORTE_DECLSPEC are not required. This commit was SVN r13825.	2007-02-27 15:45:40 +00:00
Sven Stork	d8a369936e	- Fix more symbols that should be exported. This commit was SVN r13824.	2007-02-27 15:17:17 +00:00
Sven Stork	a86deb460e	- export required symbols This commit was SVN r13810.	2007-02-27 09:43:32 +00:00
Brian Barrett	f6a5d58885	Rather than set the connect event timeout number to something big and hoping its bigger than the timeout for the connect() call, just don't register the handler by default and fall back to connect() timing out. Should give much happier performance on big clusters. This commit was SVN r13639.	2007-02-13 18:36:50 +00:00
Pak Lui	085826d94a	* Remove the code for putting the bogus exit status of the user proc. Also remove the smr set_proc_state since it's covered elsewhere. This commit was SVN r13625.	2007-02-12 23:59:27 +00:00
Brian Barrett	8b28e5b33d	Allow the OOB to connect between all MPI applications during MPI_INIT without also establishing MPI connectivity. This commit was SVN r13595.	2007-02-09 20:17:37 +00:00
Brian Barrett	262cbbc5c9	Back out r13593, which contained a change that shouldn't be committed. This commit was SVN r13594. The following SVN revision numbers were found above: r13593 --> open-mpi/ompi@81472363ea	2007-02-09 20:13:02 +00:00
Brian Barrett	81472363ea	Allow the OOB to connect between all MPI applications during MPI_INIT without also establishing MPI connectivity. This commit was SVN r13593.	2007-02-09 20:11:40 +00:00
Ralph Castain	5818a32245	Bring in a forgotten speed improvement for the TM launcher that was developed during SNL Tbird testing last year. Remove the redundant and slow calls to TM to resolve hostnames. Instead, read the host info from the PBS file during the RAS, and then just use that info in the PLS (rather than getting it again). Adjust the RMAPS mapped_node object to propagate the required launch_id info now included in the ras_node object. This provides support for those few systems that don't use nodename to launch, but instead want some id (typically an index into the array of allocated nodes). This value gets set for each node in the RAS - the RMAPS just propagates it for easy launch. This commit was SVN r13581.	2007-02-09 15:06:45 +00:00
George Bosilca	79d76b044a	ORTE_DECL everything that can be used outside the base directory. I woner why this file is called private when it's included by all PLS ... This commit was SVN r13573.	2007-02-09 03:16:19 +00:00
Ralph Castain	890e3c7981	Reset the trunk so that the odls now sets the paffinity and sched_yield params again. The sched_yield is still overridden by any user-specified setting. This change utilizes the new num_processors function. I also left the mods made to ompi_mpi_init and the bug fix for the default value of mpi_yield_when_idle. Note that the mods to mpi_init will not really take effect as the mca param will now always be set (either by user or odls). We will need those mods later, so no point in removing them now. This commit was SVN r13519.	2007-02-06 19:51:05 +00:00
Jeff Squyres	c91fcd7fbd	Fix a bunch of minor typos submitted by Bernhard Fischer. This commit was SVN r13505.	2007-02-06 12:00:30 +00:00
Ralph Castain	a8202742ba	Fix a missing function pointer - reference ticket #854 This commit was SVN r13476.	2007-02-02 23:10:14 +00:00
Ralph Castain	3daf8b341b	Fix the sched_yield problem for generic environments. We now determine and set sched_yield during mpi_init based on the following logical sequence: 1. if the user has specified sched_yield, we simply do what we are told 2. if they didn't specify anything, try to get the number of processors on this node. Note that we already now get the number of local procs in our job that are sharing this node - that now comes in through the proc callback and is stored in the ompi_proc_t structures. 3. if we can get the number of processors, compare that to the number of local procs from my job that are sharing my node. If the number of local procs exceeds the number of processors, then set sched_yield to true. If not, then be a hog and set sched_yield to false 4. if we can't get the number of processors, default to conservative behavior and set sched_yield to true. Note that I have not yet dealt with the need to dynamically adjust this setting as more processes are added via comm_spawn. So far, we are only looking within our own job. Given that we have now moved this logic to mpi_init (and away from the orteds), it isn't yet clear to me how a process will be informed about the number of procs in other jobs that are also sharing this node. Something to continue to ponder. This commit was SVN r13430.	2007-02-01 19:31:44 +00:00
Ralph Castain	c754523a14	Add cancel_operations to the pls module definition for tm This commit was SVN r13416.	2007-02-01 16:52:28 +00:00
Ralph Castain	51fb746da3	Stop overriding the yield_when_idle mca param if the user has set it This commit was SVN r13414.	2007-02-01 15:01:12 +00:00
George Bosilca	9f73335bdb	Silence the compiler. This commit was SVN r13381.	2007-01-31 04:24:56 +00:00
Jeff Squyres	8d872b195a	Refs trac:726 Tested this functionality quite a bit more and made some fixes: * Print far fewer help messages * Fix one additional deadlock upon error * Change some ORTE_LOG messages to silent (because they're not errors) * Some code got re-indented, sorry... Discussed and reviewed with Ralph. This commit was SVN r13375. The following Trac tickets were found above: Ticket 726 --> https://svn.open-mpi.org/trac/ompi/ticket/726	2007-01-30 23:03:13 +00:00
Rainer Keller	061ba05439	- Fixes uncovered with the format attribute to opal_output and opal_output_verbose This commit was SVN r13371.	2007-01-30 20:56:31 +00:00
Rainer Keller	3669e8921e	- Fix further compiler warnings regarding initialization and shadowing variables. This commit was SVN r13358.	2007-01-30 06:34:38 +00:00
George Bosilca	dea69e3c7c	Remove one of the %s. This commit was SVN r13357.	2007-01-30 03:56:48 +00:00
Rich Graham	f6c99d0207	set orte_odls_base.components_available to false if no odls components are available. Startup now works if no odls components are availble. This commit was SVN r13339.	2007-01-27 15:37:13 +00:00
Jeff Squyres	3c5c8c3c4c	Refinement of Rainer's r13227 and r13228 (worked with Rainer, Ralph, and George on these refinements): * Rename the static OBJ initializer macro to be OPAL_OBJ_STATIC_INIT(class) * Ensure that all static OBJ initializations get a refcount of 1 (doesn't ''really'' matter, since they're static, it should never get to the point where the OBJ is DESTRUCTed, but more correct nonetheless) * Add a "magic number" to the OBJ when compiling with debug support. The magic number does some rudimentary support to ensure that you're operating on a valid OBJ (and fails an assertion if you're not). Check to ensure that the memory contains the magic number when performing actions of OBJ's. Also remove the magic number when DESTRUCTing OBJs, so that if, for example, an OBJ is DESTRUCTed more than once, we'll fail the magic number assert. This commit was SVN r13338. The following SVN revision numbers were found above: r13227 --> open-mpi/ompi@96030de97b r13228 --> open-mpi/ompi@c2e9075d29	2007-01-27 13:44:03 +00:00
George Bosilca	668a2bd7ac	Remove some debug output. This commit was SVN r13323.	2007-01-26 08:09:22 +00:00
George Bosilca	bd7eebda83	Deal with the argv problem from r13321 for the Windows PLS. This commit was SVN r13322. The following SVN revision numbers were found above: r13321 --> open-mpi/ompi@b439e87f96	2007-01-26 07:21:07 +00:00
George Bosilca	b439e87f96	We have this one starting from r12059. We save a pointer to the argv[*] and then we modify the argv, forcing the reallocation of the array. With luck the saved pointer still have a meaning ... without execve return with error 14 (EFAULT). This commit was SVN r13321. The following SVN revision numbers were found above: r12059 --> open-mpi/ompi@ae79894bad	2007-01-26 07:06:52 +00:00
Rich Graham	3488b394be	fix typo in name of the cancel operation. This commit was SVN r13312.	2007-01-25 19:07:27 +00:00
Jeff Squyres	580a7a108c	Fix a compiler warning. This commit was SVN r13310.	2007-01-25 17:22:01 +00:00
Tim Prins	e199bf9b64	Refs trac:801 Fix compiler warning This commit was SVN r13308. The following Trac tickets were found above: Ticket 801 --> https://svn.open-mpi.org/trac/ompi/ticket/801	2007-01-25 16:12:05 +00:00
Ralph Castain	ab5ea61100	Bring over the rest of the ctrl-c fixes. This commit includes: 1. add a "cancel_operation" API to the pls components that allows orterun to demand that an orted operation (e.g., terminate_job) be immediately cancelled and abandoned. 2. changes the pls orted commands from blocking to non-blocking. This allows us to interrupt those operations should an orted be non-responsive. The change also adds an orte_abort_timeout that limits how long orterun will automatically wait for the orteds to respond - if the terminate command, for example, doesn't see orted response within that time, then we printout an appropriate error message and just give up. 3. modifies orterun to allow multiple ctrl-c's to simply abort the program even if the orteds have not responded 4. does some cleanup on the orte-level mca params so that their implementation looks a lot more like that of ompi - makes it easier to maintain. This change also includes the definition of an orte_abort_timeout struct and associated MCA param (can't have too many!) so you can set the time after which orterun gives up on waiting for orteds to respond This needs more testing before migrating to 1.2. This commit was SVN r13304.	2007-01-25 14:17:44 +00:00
Ralph Castain	53967bd698	Fix a memory corruption problem deep inside the registry when subscriptions/triggers are processed. The create_value function will malloc space for the pointers to keyval objects, but doesn't actually allocate space for the objects themselves. When constructing the gpr_notify_data object, we forgot to OBJ_NEW the keyval objects. Since the create_value function didn't explicitly NULL those memory locations, it just so happened that there was a non-NULL address in them....which we dutifully dumped a keyval into. This fix includes two parts: (a) we now initialize the keyval pointer locations to NULL after the malloc, and (b) we now OBJ_NEW the keyvals prior to storing info in them. BTW, in case anyone reads this and wonders why we don't just OBJ_NEW the keyvals in create_value, the reason is simply that some places in the code use static keyvals and simply assign those addresses into the value object's array. So not everyone wants to OBJ_NEW keyvals - by not forcing it here in create_value, we give the user the flexibility to do whatever they want. This commit was SVN r13300.	2007-01-25 12:54:02 +00:00
George Bosilca	dcce444ed4	When the user give a prefix that really means something. I expect to start looking for the daemons using the prefix. This commit was SVN r13297.	2007-01-25 07:35:25 +00:00
George Bosilca	3b988fcdfd	Small update the the process PLS. This commit was SVN r13293.	2007-01-25 00:17:54 +00:00
Tim Prins	4fd81b3407	Fixes trac:801 - Make it so the SLURM ras can handle different nodelist configurations - Some code cleanup and better/more informative error messages and error handling This commit was SVN r13271. The following Trac tickets were found above: Ticket 801 --> https://svn.open-mpi.org/trac/ompi/ticket/801	2007-01-24 14:45:42 +00:00
George Bosilca	9b16827049	Add ORTE_DECLSPEC, and few conversions. This commit was SVN r13268.	2007-01-24 00:52:08 +00:00
George Bosilca	1e38810c2d	Correctly close the sockets on a generic way. This commit was SVN r13254.	2007-01-23 03:17:23 +00:00
Ralph Castain	46b7df5683	Fix bproc nodename to correctly assign process-to-nodename mapping This commit was SVN r13244.	2007-01-22 19:39:00 +00:00
Jeff Squyres	6584df9262	For --prefix-like behavior, we used to modifiy environ directly and then exec the "srun..." from there. But somewhere along the line, we switched to having a copy of environ and modifying that. It looks like we forgot to update the stuff for --prefix behavior. So this commit fixes the setenv's for PATH and LD_LIBRARY_PATH to modify the environ copy (not environ itself) so that the values properly get passed down to the srun environment via execve(). This restores --prefix behavior in the SLURM pls. This commit was SVN r13239.	2007-01-22 15:50:35 +00:00
George Bosilca	1b92589179	Update the PLS process. This commit was SVN r13236.	2007-01-22 05:48:25 +00:00
Rainer Keller	c2e9075d29	- Define a OPAL_CLASS_EMPTY to be used for initialization. Similar within the dt_module for the predefined datatypes. This commit was SVN r13228.	2007-01-21 15:52:06 +00:00
Ralph Castain	b63d4ddfbf	Clean up a compiler warning under bproc This commit was SVN r13198.	2007-01-18 20:07:06 +00:00
Ralph Castain	1487e22ec8	Store the mapping mode so that it can be recovered later This commit was SVN r13197.	2007-01-18 20:00:15 +00:00
Ralph Castain	455e4ada9a	Bring the modified/updated pernode and npernode behaviors over from the openrte repository. This change enables npernode to pay attention to the total #procs to be launched, and cleans up the bynode vs. byslot mapping directives when in pernode and npernode modes. This commit was SVN r13191.	2007-01-18 17:15:19 +00:00
Brian Barrett	ffe35ef6b8	Update the Cray XT3 run-time support files to compile with latest RTE changes This commit was SVN r13172.	2007-01-17 22:47:27 +00:00
Ralph Castain	e093b5a256	Check for NULL before release, just to be safe This commit was SVN r13162.	2007-01-17 21:29:34 +00:00
Jeff Squyres	6f7adfe231	Fix for the oob base open and close functions being invoked twice by ompi_info -- once directly and once via the rml oob component. This commit was SVN r13152.	2007-01-17 15:18:13 +00:00
Jeff Squyres	3983141342	Remove this extra #if -- it wasn't necessary and was causing compiler warnings. This commit was SVN r13146.	2007-01-17 13:53:02 +00:00
Ralph Castain	f08210b3e1	Fix a double-free error This commit was SVN r13126.	2007-01-16 15:49:16 +00:00
George Bosilca	0f68bb0bc0	Add a debug flag for the ODLS process. This commit was SVN r13075.	2007-01-11 00:17:34 +00:00
George Bosilca	c8222b57eb	The Windows PLS now is able to spawn process locally. This commit was SVN r13074.	2007-01-11 00:16:58 +00:00
Rolf vandeVaart	9fd5e55b50	Need to include strings.h because that is where the rindex() function prototype lives. Without this, we get compile warnings. In addition, for 64-bit Solaris, we get a segmentation fault from orterun without this include. This commit was SVN r13065.	2007-01-10 18:44:08 +00:00
Brian Barrett	03112254e7	Increase connection timeout to 600 seconds, which should always be higher than the connect() timeout, so that we'll use that rather than our own timeout by defualt. There timeout was set low for Big Red, but causes problems for very large clusters, as there's no way to wire them up in 10 seconds most of the time. This commit was SVN r13062.	2007-01-10 04:53:21 +00:00
George Bosilca	c7da2b0a9a	Add the process PLS. It's only intended for Windows users. This commit was SVN r13051.	2007-01-09 00:19:52 +00:00
George Bosilca	77452ea8ea	Add a missing include and update the definition of orte rds proxy component. This commit was SVN r13042.	2007-01-08 22:00:01 +00:00
Brian Barrett	e130f18cc2	Fix some compiler warnings that have slipped in lately... This commit was SVN r13037.	2007-01-08 17:20:09 +00:00
Brian Barrett	a34e67d743	Remove unneeded PARAM_INIT_FILE variable in configure.params files used by components that use configure.m4 for configuration or are always built. The macro has not been needed since moving to configure types other than configure.stub Fixes trac:590 This commit was SVN r13031. The following Trac tickets were found above: Ticket 590 --> https://svn.open-mpi.org/trac/ompi/ticket/590	2007-01-08 03:44:22 +00:00
Jeff Squyres	a91c017f81	The constant name changed from ORTE_RML_NAME_ANY to ORTE_NAME_WILDCARD -- upcate the comments/documentation to match. This commit was SVN r13001.	2007-01-05 13:38:22 +00:00
Ralph Castain	3ce9b2f6cc	Remove some debugging output that mistakenly was left behind. This commit was SVN r12984.	2007-01-04 17:24:11 +00:00
Ralph Castain	6101050ea6	Remove an abstraction barrier I thought was gone long-ago. The OOB subscription really shouldn't be defined as an OMPI subscription. I know it's just a technicality, but it is time to address such things rather than just letting them continue to propagate. :-) This commit was SVN r12954.	2007-01-02 16:16:50 +00:00
Ralph Castain	90f5e3fad8	Fix a buglet in the singleton startup procedure. For purposes of minimizing the xcast message, we "strip" the descriptive info on all subscription messages. This means, though, that we have to store the process name and other info so it can be retrieved in the body of the subscription data (as opposed to in the description). This wasn't being done for singletons because they don't call the RMAPS to "map" themselves. This has now been corrected. The singleton startup will dutifully call the mapper framework so that the proper data storage locations get initialized. Unfortunately, we then had to instruct the RMAPS not to allocate a vpid range for this job - otherwise, it would make a mistake and think there were two processes in it. Hence, a change was required to RMAPS to tell it "map this job, but don't allocate a vpid range for it". This change will need to migrate across to 1.2 after it "soaks" the appropriate time. This commit was SVN r12952.	2007-01-02 16:14:44 +00:00
Rich Graham	6cb2377015	Change the allocation of the shared memory backing file. The file is allocated on a per comm_world instance, with the lowest rank in comm_world on the given host creating and initializing the file, and then notifying the remaining files via the OOB. Reviewed: Ralph Castain, Brian Barrett Addressing ticket #674. This commit was SVN r12949.	2007-01-01 02:39:02 +00:00
Brian Barrett	b5057d923e	fix compile error that crept in with orte changes. This just makes it compile -- I can't check for correctness on this platform. Someone probably wants to do that. This commit was SVN r12948.	2006-12-31 20:16:20 +00:00
Li-Ta Lo	6df4e80727	new XCPU PLS and SDS to work with libxcpu This commit was SVN r12905.	2006-12-21 00:05:36 +00:00
Brian Barrett	414a87b14c	On OS X, the lowest 8 bits of the exit status are for signals, so pushing rc (which is -1 or 4 if we hit this case) resulted in an odd error that a signal killed the proc (instead of a startup error, as is reality). Instead, use the W_EXITCODE macro (if available) to build up an exit code that has an error code for exit status, but does not make it look like the process died from a signal This commit was SVN r12890.	2006-12-18 02:30:05 +00:00
Ralph Castain	a0ef517550	Fix some errors in the bproc components that prevented compiling. Thought I had already done this, but either those changes were lost when I did the merge, or my old man's memory is fading.... Whaz-at??? :-) This commit was SVN r12874.	2006-12-15 19:40:04 +00:00
Ralph Castain	1e1d0e8a89	Set the app_num attribute into the process environment so we pick it up on the other end This commit was SVN r12868.	2006-12-15 16:43:52 +00:00
Ralph Castain	677d1260aa	cleanup nicely if we don't launch This commit was SVN r12867.	2006-12-15 14:03:53 +00:00
Ralph Castain	cbb660504c	Retain the ability to run valgrind on the bproc launcher - do not call bproc_version if "nolaunch" is specified. This commit was SVN r12866.	2006-12-15 14:01:21 +00:00
Ralph Castain	64ec238b7b	Repair support for Bproc 4 on 64-bit systems. Update the SMR framework to actually support the begin_monitoring API. Implement the get/set_node_state APIs. This commit was SVN r12864.	2006-12-15 02:34:14 +00:00
Brian Barrett	38c2e43ac2	Print out error string rather than errno for TCP-related errors, making it easier for both the user and us to debug issues with BTL and OOB issues... This commit was SVN r12852.	2006-12-14 18:20:43 +00:00
Ralph Castain	3b064a624e	For convenience, revise the orte_job_map_t object so it includes the vpid start/range values, the number of nodes, and the number of processes on each node. These values are all used in various places in the code base - we currently re-compute them multiple times. Since these values do not change and are already being computed by the RMAPS framework, we might as well just save them for re-use. This commit was SVN r12829.	2006-12-12 16:07:23 +00:00
Ralph Castain	28ce8e5e5e	Extend the mpirun options to support "--npernode N". This option tells the system to spawn N procs/node across all nodes in the allocation. If N is greater than the number of allocated slots, then the usual oversubscription logic will apply (i.e., the system will error out if oversubscription is not allowed, otherwise it will run with the sched_yield set to non-aggressive behavior). In "--npernode" operation, the "-np" command line parameter is ignored. This commit was SVN r12826.	2006-12-12 00:54:05 +00:00
Ralph Castain	8314e8dbb9	Modify the pernode option so it can accept a request for the number of processes to be launched. We now check three use-cases for pernode: 1. no -np provided - put one proc/node across all allocated nodes 2. -np N provided, N > #nodes - we print a pretty error message and exit 3. -np N provided, N <= #nodes - put one proc/node across N nodes I also added a new orte constant (ORTE_ERR_SILENT) that allows us to pass up the chain that an error was encountered, but NOT print ORTE_ERROR_LOG messages. This is intended to be used for cases where the error we encounter is NOT an orte error, but rather is one associated with incorrect user input (e.g., the preceding case 2). In such cases, there is no point in printing an ORTE_ERROR_LOG chain of messages as it isn't an orte error. This commit was SVN r12821.	2006-12-11 18:07:07 +00:00
Ralph Castain	0a5d41857a	Complete next round of message size reduction: "strip" the descriptive info from the returned values. I have now added a flag to the gpr address mode (ORTE_GPR_STRIPPED) that instructs the gpr to not include segment names or tokens in the returned gpr_value_t objects. I found only two places that were looking at the tokens: 1. the odls - we used the tokens to separately process the globals container data from everything else. In this case, I left the subscription that returned the globals data alone, but "stripped" the subscription that returned the launch data for the procs. These subscriptions have nothing to do with the xcast message. 2. the pml_base_modex - the callback function was getting process names from the returned tokens. Actually, this function was doing a very bad thing - it was assuming that the first token returned was always the process name. This is currently true, but is one of those assumptions that someone could have easily changed - and suddenly found the system inexplicably failing. I modified the function to (a) get the name sent back to us, (b) "stripped" the value structures of tokens and segment strings, and (c) correctly obtained process names from the returned values. I also reindented the heck out of the code so it was legible (at least, to my old eyes). This commit was SVN r12813.	2006-12-09 23:10:25 +00:00
Ralph Castain	58569546ed	Fix the fix to remove compiler warning - an incorrect "\" was placed in the command string. This commit was SVN r12805.	2006-12-08 04:17:38 +00:00
Sven Stork	78173a697a	Replace the test opertion "-e" with "-r" to improve the protability. Refs: #392 This commit was SVN r12790.	2006-12-07 12:14:40 +00:00
Ralph Castain	62d7826e01	Helps if we total up the correct field to get the total number of slots in the universe This commit was SVN r12789.	2006-12-07 03:17:12 +00:00
Ralph Castain	a1153fdc8f	Eliminate virtually all of the attribute_predefined data from the STG1 message. We now compute the total number of slots allocated to us and save that in the registry - the attributed_predefined then retrieves it via the STG1 message. The app_num is passed via the process_info structure, which gets the value from the ODLS in the environment. Obviously, people like bproc will have to get the app_num via another avenue...but that's a problem for another day. Several options are easily available. This commit was SVN r12788.	2006-12-07 03:11:20 +00:00
Ralph Castain	d4bd60c9fe	Restore the paffinity capability, along with all the required logic to ensure we "do the right thing" when the user gives us inaccurate information about the number of slots on a remote node. This commit was SVN r12780.	2006-12-06 15:59:34 +00:00
Ralph Castain	b1e16fffac	Add the C++ doo-hicky stuff around the odls framework definitions just in case somebody, somewhere, on some remote planet where only goats can feed needs it. This commit was SVN r12777.	2006-12-06 13:58:04 +00:00
Ralph Castain	8ca415a0c5	Remove duplicate orte_odls declaration This commit was SVN r12776.	2006-12-06 13:44:41 +00:00
Brian Barrett	6f8b366acb	Rename liborte to libopen-rte and libopal to libopen-pal per telecon today and bug #632. Refs trac:632 This commit was SVN r12762. The following Trac tickets were found above: Ticket 632 --> https://svn.open-mpi.org/trac/ompi/ticket/632	2006-12-05 18:27:24 +00:00
Tim Prins	08d5ca821f	Don't get the node architecture when useing the LoadLevleer RAS. It is slow (about a second for ~300 nodes) and we don't even use the value. This commit was SVN r12758.	2006-12-05 13:47:53 +00:00
Ralph Castain	eb941d8ae2	Fix a bug that declared a node as "oversubscribed" a little early during the mapper procedure. This only affected the mapping procedure, and only if you had set the "--no-oversubscribe" flag. Kudos to Tim Prins for finding it. This commit was SVN r12757.	2006-12-05 13:04:27 +00:00
Gleb Natapov	f0132b2499	Provide parameters in a correct order (processor/oversubscribed was swapped). This commit was SVN r12737.	2006-12-04 12:55:45 +00:00
George Bosilca	a0ed53d70b	Make the compilers happy. This commit was SVN r12729.	2006-12-03 00:19:11 +00:00
George Bosilca	3fd278c522	Make the tree compile in debug mode. This commit was SVN r12724.	2006-12-01 23:03:09 +00:00
Ralph Castain	897744cdeb	Two major changes to the runtime: 1. implement and enable the non-described buffer operations. I will send out a more detailed explanation separately. However, this mode of operation (which is now the default) significantly reduces message size during startup. If you want the described buffers, set the mca param "-mca dss_describe_buffer 1". 2. revise the xcast system to support both linear and binomial tree broadcast methods. Since we are seeing scenarios where the binomiall tree can cause problems, I have made the linear method the default. To run with the binomial tree, set the mca param "-mca oob_xcast_mode binomial". 3. add some detailed timing reports to the xcast operation. These are enabled via "-mca oob_xcast_timing 1". 4. add some more unit tests for the dss and gpr (focused on support for the non-described buffer) This commit was SVN r12722.	2006-12-01 22:30:39 +00:00
Jeff Squyres	3cf7dddd47	Fixes trac:635. Ralph identified the problem, I tracked down ''where'' the fd was being closed, and Brian figured out ''why'' (and the fix). What was happening is that a remote process was closing its stdout/stderr and therefore sending a 0-byte IOF message to mpirun. mpirun, in turn, closed the iof endpoint associated with that stream (i.e., stdout/stderr). IOF does this to handle the case where mpirun's stdin is closed -- this therefore causes the stdin on all the ORTE-started processes to have their stdin's closed as well. So the workaround here is to check that if we get a 0-byte IOF message on a sink (indicating a remote closure), and if that sink is the special stdout or stderr stream, don't actually close anything in the local process. This commit was SVN r12691. The following Trac tickets were found above: Ticket 635 --> https://svn.open-mpi.org/trac/ompi/ticket/635	2006-11-28 21:42:49 +00:00
Ralph Castain	0398c9e0c5	Correctly setup the sched_yield when launching processes via the orteds. This still doesn't adjust the yield schedule "on-the-fly" as more procs are dynamically added to a node - it just sets it when they are first launched. This commit was SVN r12683.	2006-11-28 08:27:20 +00:00
George Bosilca	8df8d86b85	Complete the functions to match the expected prototype. This commit was SVN r12680.	2006-11-28 00:44:30 +00:00
Ralph Castain	bc4e97a435	First stage in the move to a faster startup. Change the ORTE stage gate xcast into a binary tree broadcast (away from a linear broadcast). Also, removed the timing report in the gpr_proxy component that printed out the number of bytes in the compound command message as the answer was "not much" - reduces the clutter in the data. This commit was SVN r12679.	2006-11-28 00:06:25 +00:00
Ralph Castain	9bc25f0bec	Fix a potential bug in the registry where it didn't fully check a segment's name when searching for it. Will have to verify that this doesn't break other things. Bring the bproc system close to being back online.... This commit was SVN r12659.	2006-11-23 04:17:37 +00:00
Ralph Castain	deb2470ba3	Move the waitpid callback in the bproc pls after we store the daemon info. Otherwise, a short-lived app could terminate before we store the daemon info, causing mpirun to not terminate the daemons since the call to get_active_daemons would return a NULL list. This commit was SVN r12656.	2006-11-22 22:49:22 +00:00
Ralph Castain	b1ff5fe868	Move the name of the bproc common segment to the central schema location - avoids conflicts when bproc 3 components try to build This commit was SVN r12654.	2006-11-22 20:23:17 +00:00
Ralph Castain	8080034eb2	Clean up a compile issue for bproc This commit was SVN r12653.	2006-11-22 19:50:27 +00:00
Ralph Castain	428c1f14c3	Modify the bproc components to resolve the current allocation problem This commit was SVN r12652.	2006-11-22 19:10:58 +00:00
Sven Stork	dc116d4814	- Add missing mutex lock This commit was SVN r12649.	2006-11-22 13:37:58 +00:00
Ralph Castain	6fca1431f3	Back out some prior commits. These commits fixed bproc so it would run, but broke several other things (singleton comm_spawn and hostfile operations have been identified so far). Since bproc is the culprit here, let's leave bproc broken for now - I'll work on a fix for that environment that doesn't impact everythig else. This commit was SVN r12648.	2006-11-22 13:30:21 +00:00
Brian Barrett	0895f5e08d	Rename OMPI_PROCESS_NAME_{HTON, NTOH} macros to ORTE_PROCESS_NAME_{HTON, NTOH} because they are in ORTE, not OMPI. Also, remove the ORTE_PROCESS_NAME macros in iof base as they are duplicates of the ones that were in ns_types, which meant that bad things happened if you changed what an orte_process_name_t looked like. This commit was SVN r12646.	2006-11-22 03:03:21 +00:00
Ralph Castain	761d4fab25	cleanup unused variables This commit was SVN r12643.	2006-11-21 20:51:39 +00:00
Ralph Castain	a30c65ca24	Fix the allocator to make bproc happy. We were burned again by the fact that the bproc state monitor creates entries on the node segment for all the nodes in the cluster when it is opened during orte_init. As a result, the bjs allocator was never being called, and the system merrily assumed that all nodes in the cluster had been allocated to it. To fix this, I removed a test that had been inserted into the allocation procedure that checked for a non-zero node segment. This was an old artifact - the RAS components already know that they are not to overwrite any existing node segment entries (at least, bproc does - I will check the others. For now, I just want to save the bproc fix on this machine). This commit was SVN r12640.	2006-11-21 19:52:55 +00:00
Ralph Castain	050e401671	Simplify - the cellid is simply a field in the process name. Recently, we decided to just directly access it and get rid of extraneous function calls. This commit was SVN r12638.	2006-11-21 09:59:24 +00:00
Tim Prins	2afb401e39	fix some compile warnings This commit was SVN r12636.	2006-11-21 00:33:10 +00:00
Galen Shipman	c06b740220	a few more opps to get the cellid... This fixes a compilation error on bproc machines.. This commit was SVN r12631.	2006-11-20 19:45:01 +00:00
Ralph Castain	33affed09c	Bring ortehalt to a preliminary capability. It will corectly order a persistent daemon to exit cleanly. Need to now interface it to orterun, clean up a few things here and there This commit was SVN r12626.	2006-11-18 04:47:51 +00:00
Ralph Castain	ea1e0d34c8	Fix the SLURM PLS and SDS components so they correctly communicate the name of the node. This repairs a number of comm_spawn issues seen on Odin and other SLURM machines. This commit was SVN r12621.	2006-11-17 20:51:03 +00:00
Ralph Castain	3c5a2cd17b	Cleanup a few warnings for unused variables This commit was SVN r12620.	2006-11-17 19:32:49 +00:00
Ralph Castain	f771cc4fbd	Modify the reuse daemons procedure so we only generate the add_local_procs message once. Revise the display-map-at-launch option so the RMAPS framework takes responsibility for implementation of that option. Modify the RMAPS framework so we eliminate communicating a map to a backend node when certain attributes are set. The proxy functions are now implemented in the base, and a check made for HNP/non-HNP operation made in the map_jobs function prior to execution. This commit was SVN r12619.	2006-11-17 19:06:10 +00:00
Ralph Castain	ca5b4358fa	Need to revise the display-map-at-launch option so it is active not only for the initial launch, but applies to any subsequent comm_spawn events too. Add placeholders for the new orte tools. These don't actually do anything yet - in fact, I have set the .ompi_ignore so that you won't compile them (I have set a .ompi_unignore for me). Please let me know if you encounter any trouble with this - the ompi_ignore's should protect everyone. This commit was SVN r12616.	2006-11-17 02:58:46 +00:00
Ralph Castain	5ddcb8a652	Ensure the orted kills all local procs when exiting. Add a little clarity to some of the debugging output This commit was SVN r12615.	2006-11-16 21:15:25 +00:00
Ralph Castain	c1813e5c5a	Extend the daemon reuse functionality to most of the other environments. Note that Bproc won't support this operation, so we just ignore the --reuse-daemons directive. I'm afraid I don't understand the POE and XGrid environments well enough to attempt the necessary modifications. Also, please note that XGrid support has been broken on the trunk. I don't understand the code syntax well enough to make the required changes to that PLS component, so it won't compile at the moment. I'm hoping Brian has a few minutes to fix it after SC. This commit was SVN r12614.	2006-11-16 15:11:45 +00:00
Ralph Castain	a17e27dfd5	Sign of old age.....fix some compiler complaints This commit was SVN r12611.	2006-11-15 22:28:14 +00:00
Ralph Castain	f7fc19a2ca	Create the ability to re-use existing daemons. Included in the commit: 1. new functionality in the pls base to check for reusable daemons and launch upon them 2. an extension of the odls API to allow each odls component to build a notify message with the "correct" data in it for adding processes to the local daemon. This means that the odls now opens components on the HNP as well as on daemons - but that's the price of allowing so much flexibility. Only the default odls has this functionality enabled - the others just return NOT_IMPLEMENTED 3. addition of a new command line option "--reuse-daemons" to orterun. The default, for now, is to NOT reuse daemons. Once we have more time to test this capability, we may choose to reverse the default. For one thing, we probably want to investigate the tradeoffs in start time for comm_spawn'd processes that reuse daemons versus launch their own. On some systems, though, having another daemon show up can cause problems - so they may want to set the default as "reuse". This is ONLY enabled for rsh launch, at the moment. The code needing to be added to each launcher is about three lines long, so I'll be doing that as I get access to machines I can test it on. This commit was SVN r12608.	2006-11-15 21:12:27 +00:00
Andrew Friedley	3f550c3a83	missing a semicolon... This commit was SVN r12607.	2006-11-15 20:35:27 +00:00
Ralph Castain	437f2b044d	Modify the orted command communication system in two ways: 1. use non-blocking sends to transmit commands (this was actually done in a prior commit) 2. have an "ack" message sent back from the orted when it completes the command The latter item is the new one here. With my prior commit, it was possible for the HNP to move on to other things before the orted had completed its command. This caused the HNP to occassionally exit before the orted, thus generating "lost connection" errors. With this change, we retain the parallel nature of the command communications, but still hold the HNP at that point until the orteds are done. Best of both worlds. This commit was SVN r12605.	2006-11-15 15:09:28 +00:00
Ralph Castain	22a473a831	Fix a compiler complaint for too many arguments to a printf command - thanks to Tim Prins for pointing it out to me. This commit was SVN r12604.	2006-11-15 14:52:48 +00:00
Ralph Castain	6d6cebb4a7	Bring over the update to terminate orteds that are generated by a dynamic spawn such as comm_spawn. This introduces the concept of a job "family" - i.e., jobs that have a parent/child relationship. Comm_spawn'ed jobs have a parent (the one that spawned them). We track that relationship throughout the lineage - i.e., if a comm_spawned job in turn calls comm_spawn, then it has a parent (the one that spawned it) and a "root" job (the original job that started things). Accordingly, there are new APIs to the name service to support the ability to get a job's parent, root, immediate children, and all its descendants. In addition, the terminate_job, terminate_orted, and signal_job APIs for the PLS have been modified to accept attributes that define the extent of their actions. For example, doing a "terminate_job" with an attribute of ORTE_NS_INCLUDE_DESCENDANTS will terminate the given jobid AND all jobs that descended from it. I have tested this capability on a MacBook under rsh, Odin under SLURM, and LANL's Flash (bproc). It worked successfully on non-MPI jobs (both simple and including a spawn), and MPI jobs (again, both simple and with a spawn). This commit was SVN r12597.	2006-11-14 19:34:59 +00:00
Ralph Castain	f95e20e2e1	Add another test program - an MPI app that just spins. This supports testing of system response to signal-terminated processes. Add some debugger output to the ODLS default component. Modify the orted command communication system so that it is done via non-blocking sends. This removes the linearity of the transmission and improves the response time. This commit was SVN r12585.	2006-11-13 21:51:34 +00:00
Tim Prins	39bc652899	Refs trac:612 Make it so if -np was not passed and -pernode was, we map bynode This commit was SVN r12580. The following Trac tickets were found above: Ticket 612 --> https://svn.open-mpi.org/trac/ompi/ticket/612	2006-11-13 19:13:21 +00:00
Ralph Castain	4636125e2d	Modify the RMGR components to allow job setup with a given jobid, and add another attribute so that we can setup triggers without launching. Add some debugging output to the ODLS default module, and the orted. Remove the nodename data from the ODLS info report - that info is already stored in the registry by the RMAPS framework upon completing the mapping procedure. Add another test program that does an ORTE-only dynamic spawn (gasp!). Looks just like comm_spawn - just no MPI involved. Modify the ODLS to release the processor when we "kill" local procs in a more scalable fashion. It previously had a sleep in it that Jeff's prior commit removed. However, he introduced some Windows code into the non-Windows component (protected by "if"s, but unnecessary). This is a more general solution he proposed - included here so I could get things to compile properly. This commit was SVN r12579.	2006-11-13 18:51:18 +00:00
Jeff Squyres	bfdf801487	Replace a "sleep(1)" with a yield so that the orted can reap processes much faster. This commit was SVN r12575.	2006-11-13 12:45:03 +00:00
Ralph Castain	4e50cdae52	This commit accomplishes two things: 1. Fix the "hang" condition when an application isn't found. It turned out that the ODLS had some difficulty with the process actually not having been started - hence, it never called the waitpid callback. As a result, the "terminated" trigger didn't fire, and so mpirun didn't wake up. With this change, the HNP's errmgr forces the issue by causing the trigger to fire itself when an abort condition occurs. 2. Shift the recording of the pid and the nodename from mpi_init to the orted launcher. This allows programs such as Eclipse PTP to get the pids even for non-MPI applications. In the case of bproc, the pls handles this chore since we don't use orteds in that system. This commit was SVN r12558.	2006-11-11 04:03:45 +00:00
Sven Stork	9ba4c4a7ee	- Add support for "sh" handling. Instead of detecting of bash we now check for bourne shell, because bourne shell is the smallest common divisor for bash/ksh/sh. - Make some shell expressions sh compatible This commit was SVN r12509.	2006-11-09 10:16:45 +00:00
Ralph Castain	ea77beca29	cleanup the TM modules in prep for T-bird tests. The TM RAS will now report time required to resolve hostnames This commit was SVN r12449.	2006-11-06 20:56:18 +00:00
Ralph Castain	884caeb2c7	Add timing tests for the TM ras This commit was SVN r12445.	2006-11-06 18:41:22 +00:00
Galen Shipman	68d9922f44	enable/disable connection sleep in oob_tcp.c via mca param.. on by default.. This commit was SVN r12444.	2006-11-06 18:00:46 +00:00
Ralph Castain	d182ae7472	Clean up a few compiler warnings courtesy of Jeff This commit was SVN r12430.	2006-11-03 20:45:22 +00:00
Ralph Castain	c77f6c605e	Update timing reports: 1. Remove timing of xcast from mpi_init 2. Add timing report from oob_xcast on how long it took to send the message This commit was SVN r12428.	2006-11-03 18:55:05 +00:00
Ralph Castain	761c8beeb7	Add output of the compound command size to the timing reports This commit was SVN r12423.	2006-11-03 16:11:37 +00:00
Ralph Castain	60e27c77e7	Add some additional timing reporting: 1. Added reporting points around the xcasts in MPI_Init. Note that these times will include time spent waiting for a trigger to fire, which is why the times between stage gates did NOT include these times initially. The inter-stage-gate times still do NOT include the xcast time - the xcast time is reported separately. 2. Added the process vpid on the MPI_Init timing reports for clarity. 3. Added a report from the xcast function on the HNP that outputs the number of bytes in the message being sent to the processes. This commit was SVN r12422.	2006-11-03 16:04:40 +00:00
Ralph Castain	2472f76d89	Make sure the print_map doesn't get called if we don't do the map step in rmgr.spawn This commit was SVN r12404.	2006-11-02 05:52:24 +00:00
Ralph Castain	dc4de57253	Per request from the Eclipse team, change the attributes that control the "flow" through the rmgr.spawn procedure from "stop after this step" to "do this step". This allows them to use the spawn procedure to complete the launch one step at a time so they can let the user see the results after each step. Shouldn't affect anyone else. This commit was SVN r12403.	2006-11-02 05:15:24 +00:00
Ralph Castain	233dac8bba	Enable the new attributes for remote operations such as comm_spawn This commit was SVN r12382.	2006-10-31 23:52:25 +00:00
Ralph Castain	30de73a712	Add a few attributes that are helpful for folks doing things like Eclipse. Also add yet another command-line option to orterun to support one of the new attributes. These include: 1. ORTE_RMAPS_DISPLAY_AT_LAUNCH: pretty-prints out the process map right before we launch so you can see where everyone is going. This is settable via the command line option "--display-map-at-launch" 2. ORTE_RMGR_STOP_AFTER_SETUP: just setup the job and then return from the spawn command. 3. ORTE_RMGR_STOP_AFTER_ALLOC: return from the rmgr.spawn call after allocating the job 4. ORTE_RMGR_STOP_AFTER_MAP: return from the rmgr.spawn call after mapping the job. This gives folks a chance to retrieve and graphically display the map, let the user edit it, and store the results. They can then call "launch" on their own and the system will use the revised map. Enjoy! My personal favorite is the first one - helps with debugging. This commit was SVN r12379.	2006-10-31 22:16:51 +00:00
Ralph Castain	17c71f8d2a	Fix ticket #545 Setup subscriptions to correctly return the MPI_APPNUM attribute. Fix an unreported bug that was found. The universe size was incorrectly defined in the attributes code. As coded, it looked for size_t values and based its size computation on those numbers. Unfortunately, the node_slots value had been changed to an orte_std_cntr_t awhile back! So the universe size was never updated. Update the hello_nodename test to check for MPI_APPNUM. Add a definition to ns_types for ORTE_PROC_MY_NAME - just a shortcut for orte_process_info.my_name. Brought over from ORTE 2.0 as it will be used extensively there. This commit was SVN r12377.	2006-10-31 21:29:07 +00:00
Tim Prins	894b220fbb	Fixes trac:487 Give a more intelligible error message when someone passes -nolocal and the only available node is the local node. This commit was SVN r12325. The following Trac tickets were found above: Ticket 487 --> https://svn.open-mpi.org/trac/ompi/ticket/487	2006-10-26 21:46:18 +00:00
Ralph Castain	36d4511143	Bring the timing instrumentation to the trunk. If you want to look at our launch and MPI process startup times, you can do so with two MCA params: OMPI_MCA_orte_timing: set it to anything non-zero and you will get the launch time for different steps in the job launch procedure. The degree of detail depends on the launch environment. rsh will provide you with the average, min, and max launch time for the daemons. SLURM block launches the daemon, so you only get the time to launch the daemons and the total time to launch the job. Ditto for bproc. TM looks more like rsh. Only those four environments are currently supported - anyone interested in extending this capability to other environs is welcome to do so. In all cases, you also get the time to setup the job for launch. OMPI_MCA_ompi_timing: set it to anything non-zero and you will get the time for mpi_init to reach the compound registry command, the time to execute that command, the time to go from our stage1 barrier to the stage2 barrier, and the time to go from the stage2 barrier to the end of mpi_init. This will be output for each process, so you'll have to compile any statistics on your own. Note: if someone develops a nice parser to do so, it would be really appreciated if you could/would share! This commit was SVN r12302.	2006-10-25 15:27:47 +00:00
Brian Barrett	d6ff14ed61	Hand-pack the connection information for each peer rather than just packing a sockaddr_in, as there are some endianness and padding issues with sending a sockaddr_in. Note that the sin_port and sin_addr are already in network byte order, which is why we pack them as a byte string. Refs trac:493 This commit was SVN r12301. The following Trac tickets were found above: Ticket 493 --> https://svn.open-mpi.org/trac/ompi/ticket/493	2006-10-25 15:09:30 +00:00
Ralph Castain	601443e690	Fix the other end of the attribute store/retrieve to correctly handle attributes with no value This commit was SVN r12290.	2006-10-25 01:41:58 +00:00
Ralph Castain	46166a9c77	Fix a bug that would cause a segfault when someone specified an option that had no corresponding value This commit was SVN r12288.	2006-10-24 22:06:18 +00:00
Ralph Castain	a7acd22e47	Fix a fix so we see the correct MCA param This commit was SVN r12282.	2006-10-24 17:24:09 +00:00
George Bosilca	ed3caa9fc7	mca_base_param_reg_int require a pointer to a component not a name. This commit was SVN r12281.	2006-10-24 16:41:51 +00:00
George Bosilca	7dec7731ce	Instead of size_t use orte_std_cntr_t. Remove all warnings. This commit was SVN r12280.	2006-10-24 16:40:49 +00:00
Ralph Castain	8636ac6a4d	Fix ticket 353 - print out a nice message that the combination of debug-daemons and num_concurrent in the pls rsh launcher will cause deadlock and exit This commit was SVN r12279.	2006-10-24 15:59:02 +00:00
Tim Prins	cb622db7c9	Fixes trac:352 Only close off stdout/stderr from the daemons if we are not debugging the slurm pls and --debug-daemons was not passed. This commit was SVN r12276. The following Trac tickets were found above: Ticket 352 --> https://svn.open-mpi.org/trac/ompi/ticket/352	2006-10-24 13:05:13 +00:00
Tim Prins	93d61d01fb	Fix for a problem on SLURM we have neen having since r12243 where mpirun would hang after the process had finished. It turns out that we were always reporting the name of the daemon wrong, but we simply never noticed as we never used it, until r12243. This makes it so we report the name of the daemon correctly. This commit was SVN r12274. The following SVN revision numbers were found above: r12243 --> open-mpi/ompi@153e38ffc9	2006-10-24 01:41:28 +00:00
Ralph Castain	7a77ef0ae3	Given the amount of pain singletons cause, one can't help but wonder if it REALLY was that much trouble for people to type "mpirun -n 1 foo"....sigh. Get the ordering right so that a singleton can start. Protect the rmgr copy app_context function from NULL fields Tell the mapper it is okay for there not to be a pre-existing mapping plan for a parent when dynamically spawning processes This commit was SVN r12257.	2006-10-23 15:15:45 +00:00
George Bosilca	2a863df0a5	Newline is required by some compilers at the end of a file. This commit was SVN r12244.	2006-10-21 05:56:04 +00:00
Ralph Castain	153e38ffc9	Lesson to be learned: if you send an ack to a recv'd command, better not send it to the same tag it came from - at least, not if there is a persistent recv on that tag! Fix the persistent daemon problem where it was exiting when a job completed. Problem was that the persistent daemon would order the job daemons to exit. They would then send an 'ack' back to the persistent daemon - but the ack consisted of an echo of the "exit" command, which was recv'd by the wrong listener who treated it as a properly sent cmd....and exited. This commit was SVN r12243.	2006-10-21 02:53:19 +00:00
Ralph Castain	ab7bbb80a5	Teach the mapper to correctly handle the unbalanced --host scenario. We now map in a more expected fashion. This commit was SVN r12240.	2006-10-20 20:48:24 +00:00
George Bosilca	548b94e4e1	Add a missing ORTE_DECLSPEC. This commit was SVN r12236.	2006-10-20 19:37:21 +00:00
Tim Prins	28bf4d85ab	A couple of small fixes: - It is possible to leave a byslot/bynode routine and have cur_node_item be NULL, so check for that. - After we do an allocation where the user has provided a map (i.e. with --host), cur_node_item is pointing into the map list, not the global list. Change it to point into the global list. This commit was SVN r12232.	2006-10-20 19:00:17 +00:00
Ralph Castain	955d11fa7b	The bookmark now respects slot assignments a little better. It will not oversubscribe the first node, but will take only what is available there before moving on. See the comment in orte/mca/rmaps/round_robin/rmaps_rr.c if you want the details... :-) This commit was SVN r12230.	2006-10-20 18:24:14 +00:00
Ralph Castain	ec0bb9ffda	Fix the bookmark system - we now have children being correctly spawned where they should! Also, I am no longer seeing any issue with the child job spawning its own daemons - this appears to be fixed. We still don't reuse the existing daemons, however, but that will come. This commit was SVN r12229.	2006-10-20 18:05:16 +00:00
Ralph Castain	c07d4e2510	Cleaner rendition now extended to other environments. Remove MCA params for backend procs that can cause trouble. Specifically, any directives on the selection of components for RDS, RAS, RMAPS, PLS, and RMGR can be bad mojo on the backend. This patch will cause a problem for cnos, however, as there we want to specifically tell the backends to be "null". I'm working on that issue. This commit was SVN r12225.	2006-10-20 16:50:13 +00:00
Ralph Castain	02efd07b60	Fix the MCA param passing issue, at least for rsh at the moment. I will clean this up and move it to the other environments once I shift back to a local computer. This commit was SVN r12224.	2006-10-20 15:27:29 +00:00
Ralph Castain	b07a6b1d7a	Fix a major typo that caused remote launch to crash - had something inside the wrong brace This commit was SVN r12221.	2006-10-20 14:30:23 +00:00
George Bosilca	2aa3e51223	Nothing relevant. Only a set of castings to have a clean compile on Windows. The cl.exe compiler is pretty good at complaining about any kind of non explicit cast. This commit was SVN r12207.	2006-10-20 02:25:50 +00:00
George Bosilca	7982a23bde	ORTE_DECLSPEC should be ... This commit was SVN r12206.	2006-10-20 02:23:54 +00:00
George Bosilca	5a939e21b2	Populate the file with ORTE_DECLSPEC declarations. This commit was SVN r12205.	2006-10-20 02:19:30 +00:00
Tim Prins	ade94b523b	Fixed a number of issues related to resource allocation: - Simplified the logic of the ras modules by moving the attribute handling into the base allocation function. This allows us to decide how to allocate based on the situation, and solves some of the allocation problems we were having with comm_spawn. - moved the proxy component into the base. This was done because we always want to call the proxy functions if we are not on a HNP regardless of the attributes passed. - Got rid of the hostfile component. What little logic was in it was moved into the base to deal with other circumstances. The hostfile information is currently being propagated into the registry by the RDS, so we just use what is already in the registry. - renamed some slurm function so that they have the proper prefix. Not strictly necessary as they were static, but it makes debugging much easier. - fixed a buglet in the round_robin rmaps where we would return an error when really no error occured. I tried to make proper corrections to all the ras modules, but I cannot test all of them. This commit was SVN r12202.	2006-10-19 23:33:51 +00:00
Ralph Castain	ab196c3121	Okay, this fixes the problem of MCA params spreading too far. Sorry for the multiple corrections. This commit was SVN r12201.	2006-10-19 22:51:02 +00:00
George Bosilca	c788f0dd51	Apparently, the MCA params are being loaded into the environment params of the individual app_contexts as well. Clear them of "bad" ones. This commit was SVN r12198.	2006-10-19 21:19:08 +00:00
Ralph Castain	d0d3d7fd41	Apparently, the MCA params are being loaded into the environment params of the individual app_contexts as well. Clear them of "bad" ones. This commit was SVN r12197.	2006-10-19 20:49:35 +00:00
Ralph Castain	263f4379e8	Clean up an error in the mapper that caused "-hosts" to bomb. Update the mapper so it correctly points to the next node to be used if we are mapping by slots. As it was, if we had an app_context that used only one slot on a node, the next app_context would start on the next node - leaving a blank slot in-between. This commit was SVN r12193.	2006-10-19 18:57:29 +00:00
Tim Prins	81d400ddfd	break when they are equal, not not equal This commit was SVN r12182.	2006-10-18 21:47:01 +00:00
Tim Prins	ab964d096a	Need a terminating NULL... This commit was SVN r12180.	2006-10-18 20:52:31 +00:00
Ralph Castain	d0eb7d7216	Complete the attribute management functions. Modify the mapper to better bookmark its stopping place each time, and to pick up the next time from there. This needs to be validated on a multi-node system. Fix a major memory corruption problem in the registry put/get functions that was doing multiple free's. Not sure how valgrind missed this one, though it only occurred in specific circumstances (such as comm_spawn). This commit was SVN r12179.	2006-10-18 20:02:16 +00:00
Ralph Castain	f4a458532b	This doesn't totally resolve the comm_spawn problem, but it helps a little. I'll continue working on it and hope to resolve it completely shortly. The issue primarily centers on where to start mapping the child job's processes, and how to deal with oversubscription that might result. At the moment, I am trying to resolve the first issue first (hey, that even sounds right!). This change does a couple of things: 1. Since the USE_PARENT_ALLOC attribute is a directive about regarding allocation of resources to a job, it more properly should be an attribute of the RAS. Change the name to reflect that and move the attribute define to the ras_types.h file. 2. Add the attributes list to the RMAPS map_job interface. This provides us with the desired flexibility to dynamically specify directives for mapping. The system will - in the absence of any attribute-based directive - default to the values provided in the MCA parameters (either from environment or command-line interface). This commit was SVN r12164.	2006-10-18 14:01:44 +00:00
George Bosilca	1d69375dd7	Somehow, somewhere we have to initialize rc ... This commit was SVN r12156.	2006-10-17 22:32:21 +00:00
Ralph Castain	0c0fe022ff	This is a first cut at fixing the problem of comm_spawn children being mapped onto the same nodes as their parents. I am not convinced the behavior implemented here is the long-term right one, but hopefully it will help alleviate the situation for now. In this implementation, we begin mapping on the first node that has at least one slot available as measured by the slots_inuse versus the soft limit. If none of the nodes meet that criterion, we just start at the beginning of the node list since we are oversubscribed anyway. Note that we ignore this logic if the user specifies a mapping - then it's just "user beware". The real root cause of the problem is that we don't adjust sched_yield as we add processes onto a node. Hence, the node becomes oversubscribed and performance goes into the toilet. What we REALLY need to do to solve the problem is: (a) modify the PLS components so they reuse the existing daemons, (b) create a way to tell a running process to adjust its sched_yield, and (c) modify the ODLS components to update the sched_yield on a process per the new method Until we do that, we will continue to have this problem - all this fix (and any subsequent one that focuses solely on the mapper) does is hopefully make it happen less often. This commit was SVN r12145.	2006-10-17 19:35:00 +00:00
Tim Prins	720eb88cad	Make no-op function match new interface. This commit was SVN r12142.	2006-10-17 17:34:06 +00:00
Tim Prins	8b0170148e	Add some missing headers. This commit was SVN r12141.	2006-10-17 17:28:02 +00:00
Tim Prins	5d31332f97	Goodbye poe, long live LoadLeveler... This commit was SVN r12140.	2006-10-17 17:07:48 +00:00
Ralph Castain	13227e36ab	This commit looks a lot bigger than it is, so relax :-) Fix the problem observed by multiple people that comm_spawned children were (once again) being mapped onto the same nodes as their parents. This was caused by going through the RAS a second time, thus overwriting the mapper's bookkeeping that told RMAPS where it had left off. To solve this - and to continue moving forward on the ORTE development - we introduce the concept of attributes to control the behavior of the RM frameworks. I defined the attributes and a list of attributes as new ORTE data types to make it easier for people to pass them around (since they are now fundamental to the system, and therefore we will be packing and unpacking them frequently). Thus, all the functions to manipulate attributes can be implemented and debugged in one place. I used those capabilities in two places: 1. Added an attribute list to the rmgr.spawn interface. 2. Added an attribute list to the ras.allocate interface. At the moment, the only attribute I modified the various RAS components to recognize is the USE_PARENT_ALLOCATION one (as defined in rmgr_types.h). So the RAS components now know how to reuse an allocation. I have debugged this under rsh, but it now needs to be tested on a wider set of platforms. This commit was SVN r12138.	2006-10-17 16:06:17 +00:00
Rainer Keller	3f88937081	- Error logging is really not yet enabled. - Correct the error log for orte_errmgr_base_select - Spelling fixes This commit was SVN r12135.	2006-10-17 09:11:20 +00:00
Ralph Castain	16e52c5784	Fix a non-compliance issue regarding hostfiles as reported by Sun. The man page states that an entry that specifies slots_max but does not specify "slots" will have the soft limit default to the hard limit. The hostfile implementation, however, defaulted the soft limit to 1. This fix changes that behavior to conform to the man page. This commit was SVN r12129.	2006-10-17 00:43:12 +00:00
Ralph Castain	3f55d6897a	Remove the memory debugging options. Fix what appears to be a typo in a help file. This commit was SVN r12107.	2006-10-12 00:44:48 +00:00
Brian Barrett	fce5130333	Delay opening the listen socket until module init, so that we can have the seed value have something set to true. Allow selection of the listen type to thread if (and only if) the process is the HNP... This commit was SVN r12105.	2006-10-11 21:29:29 +00:00
Brian Barrett	29c91cf2f3	* Fix issue in odls_bproc where we were using vpid instead of the number of processes launched locally for the stdio file names. This was causing the expected files to not exist and bproc_vexecmove_io to fail. * Clean up a bunch of debugging output in the bproc pls This commit was SVN r12102.	2006-10-11 20:34:12 +00:00
Ralph Castain	f91a95b3fe	Fix the bug that caused mpirun to hang when a remote executable wasn't found using the rsh launcher. Will now test on a remote node This commit was SVN r12095.	2006-10-11 18:43:13 +00:00
Ralph Castain	2da8245be0	Correctly propagate no-daemonize This commit was SVN r12093.	2006-10-11 17:53:17 +00:00
Ralph Castain	e5877cc459	Add the proper valgrind params This commit was SVN r12092.	2006-10-11 17:48:41 +00:00
George Bosilca	b56636c855	orte_pls belong to the PLS framework, therefore it should only be defined in pls.h. This commit was SVN r12089.	2006-10-11 17:12:22 +00:00
Ralph Castain	27e305347c	Add a couple of options to orterun that support debugging of daemons for memory corruption. Ensure that the environment provided to local application processes isn't "polluted" by the orteds This commit was SVN r12087.	2006-10-11 15:18:57 +00:00
Brian Barrett	f5b8f1f2f0	Work around Automake not knowing how to properly configure libtool to build Objective C libraries Refs trac:483 This commit was SVN r12080. The following Trac tickets were found above: Ticket 483 --> https://svn.open-mpi.org/trac/ompi/ticket/483	2006-10-10 20:14:26 +00:00
Ralph Castain	699ffcf359	Restore the "bynode" mapping functionality - accidentally deleted setting of parameter This commit was SVN r12078.	2006-10-10 19:41:22 +00:00
George Bosilca	7dadc1832d	Correctly export the required functions. They are defined in a private file, but they are completely public. This commit was SVN r12070.	2006-10-10 04:54:51 +00:00
Ralph Castain	0e9dc590b7	Fix typo that didn't make it over from testing on vogon This commit was SVN r12068.	2006-10-09 20:37:39 +00:00
Ralph Castain	cebdc51762	Remove a debugging output This commit was SVN r12066.	2006-10-09 02:10:52 +00:00
Ralph Castain	2e09128337	Many thanks to Jeff for tracking down the typo causing the orte_job_map_t destuctor to fail!! Restore the OBJ_RELEASE calls to cleanup map objects. This commit was SVN r12064.	2006-10-07 22:44:00 +00:00
Ralph Castain	98dd57b70e	Add a new option to launch "pernode" - launches one process/node across all available nodes. The other options also work correctly: "-bynode" with no -np will launch on all slots, mapped on a per-node basis. This commit was SVN r12063.	2006-10-07 19:50:12 +00:00
Jeff Squyres	efe28d62e9	Fix some compiler errors. I have not checked this for correctness; but it does now compile. This commit was SVN r12062.	2006-10-07 19:10:56 +00:00
Ralph Castain	ae79894bad	Bring the map fixes into the main trunk. This should fix several problems, including the multiple app_context issue. I have tested on rsh, slurm, bproc, and tm. Bproc continues to have a problem (will be asking for help there). Gridengine compiles but I cannot test (believe it likely will run). Poe and xgrid compile to the extent they can without the proper include files. This commit was SVN r12059.	2006-10-07 15:45:24 +00:00
Ralph Castain	82a023c731	Fix a typo that caused a segfault if a caller requested that we abort an array of procs (i.e., MPI_Abort when it specifies the other procs to be aborted). Should hopefully address the recent problem seen with the BLACS AUX test as discussed on the user mailing list. This commit was SVN r12055.	2006-10-07 01:58:11 +00:00
Ralph Castain	ee0df85ece	Add a stupid, useless const since someone put it in the odls.h file. This commit was SVN r12054.	2006-10-07 01:42:23 +00:00
George Bosilca	cda46efd2a	Some missing DECLSPEC This commit was SVN r12047.	2006-10-06 15:21:52 +00:00
George Bosilca	017af37291	Keep only the useful _DECLSPEC and _DECLSPEC some globals. This commit was SVN r12042.	2006-10-06 07:24:34 +00:00
George Bosilca	b7579b09c7	Correctly handle the ORTE_DECLSPEC attribute. This commit was SVN r12041.	2006-10-06 07:08:17 +00:00
George Bosilca	422ce1d3f8	If we are in the ODLS framework then the types should be called odls not pls. This commit was SVN r12040.	2006-10-06 07:05:58 +00:00
George Bosilca	7fed79434e	Windows is now able to create local processes. This commit was SVN r12039.	2006-10-06 07:04:43 +00:00
George Bosilca	fa01b9b9aa	Last step for the name reversion (from windows back to process). This commit was SVN r12014.	2006-10-05 06:36:11 +00:00
George Bosilca	b7a793a6db	Rename all the files in the process directory. This commit was SVN r12013.	2006-10-05 06:34:30 +00:00
George Bosilca	fee0909815	Rename the Windows component. This commit was SVN r12012.	2006-10-05 06:32:57 +00:00
George Bosilca	c79c436c8d	Cleanups. Remove all __WINDOWS__ checks as this module will never get compiled on Windows. This commit was SVN r12011.	2006-10-05 06:17:30 +00:00
George Bosilca	dbe7f8ac32	Always return bool. This commit was SVN r12009.	2006-10-05 05:45:18 +00:00
George Bosilca	d1e884fbf5	Make sure we always return a bool (true/false). This commit was SVN r12002.	2006-10-05 05:27:46 +00:00
George Bosilca	3a34f9340e	If the enum is defined inside the struct it will has a scope. We don't really need that. This commit was SVN r12001.	2006-10-05 05:27:04 +00:00
George Bosilca	090b8a9098	opal_list_is_empty return true or false ... This commit was SVN r12000.	2006-10-05 05:26:08 +00:00
George Bosilca	03083cc1f6	Don't release the values[0] before it get initialized. This commit was SVN r11999.	2006-10-05 05:25:18 +00:00
George Bosilca	fd76e56279	One protection against C++ compilers is more than enough. This commit was SVN r11998.	2006-10-05 05:24:43 +00:00
George Bosilca	ad5810e33f	ORTE_DECLSPEC what needs to be ORTE_DECLSPES. This commit was SVN r11997.	2006-10-05 05:22:22 +00:00
Ralph Castain	faf3a558e6	Missing CR at end of file This commit was SVN r11959.	2006-10-03 18:17:52 +00:00
Ralph Castain	cd7d87aa7b	Define the map data types for dss compatibility. Setup to debug bproc This commit was SVN r11955.	2006-10-03 17:40:00 +00:00
Ralph Castain	4e39878944	Add a "dump" capability to the DSS so one can display a single data value to an output stream. Add some comments to the map type def in prep for building its data type support. This commit was SVN r11947.	2006-10-03 08:40:35 +00:00
Ralph Castain	99f2986db7	Bring comm_spawn back online. Shift the trigger hosting responsibilities to the HNP. We still have an issue with the io forwarding going through the spawning process, but that will be dealt with at a future time. This commit was SVN r11943.	2006-10-03 02:07:58 +00:00
Ralph Castain	b269e4da9b	Add missing functionaltiy to the ns replica to support remote get_job_peers requests. Add trace commands to help try and track down remaining problem with comm_spawn. This commit was SVN r11939.	2006-10-02 19:44:35 +00:00
Ralph Castain	9eb14425b7	The last of the debug messages that keep hiding. My apologies. This commit was SVN r11937.	2006-10-02 18:43:32 +00:00
Ralph Castain	3fd67a038f	Bring comm_spawn and persistent operations online. Still some minor problems, though - so don't use them yet, please. This won't affect anything except those two scenarios. This commit was SVN r11936.	2006-10-02 18:29:15 +00:00
Ralph Castain	12328395ae	Missed a couple of debug statements This commit was SVN r11935.	2006-10-02 15:46:41 +00:00
Ralph Castain	7494a7a83f	Clean out some debugging statements that were inadvertently left in the commit This commit was SVN r11933.	2006-10-02 15:03:18 +00:00
Ralph Castain	559b9b0ae8	Continue beating on comm_spawn. Setup to debug bproc. This commit was SVN r11932.	2006-10-02 14:58:22 +00:00
Ralph Castain	65593cd67e	Fix a few Cyrador warnings This commit was SVN r11930.	2006-10-02 13:00:32 +00:00
Brian Barrett	8f7ab1c584	num_procs can be zero if something went partly wrong before. This will cause a math exception on some platforms, so don't let that happen. This commit was SVN r11929.	2006-10-02 01:27:22 +00:00
Ralph Castain	121f834776	Continue bringing comm_spawn back online. Ensure all RM frameworks post their HNP receives. Fix the rmgr proxy component. Still need some work on the proxy component, and on job termination for persistent daemon case. This commit was SVN r11928.	2006-10-02 00:46:31 +00:00
Brian Barrett	e464adcd51	Need to tell the daemon how many procs it will start (always 1, because of the way we do the fake mapping thing... This commit was SVN r11924.	2006-10-01 23:25:22 +00:00
Brian Barrett	95ba51fbd4	* Clean up debugging output so that it's useful * Error message in an NSError object is localizedDescription, not localizedErrorReason. The latter is a decription of how the error can occur, which is usually nothing in XGrid frameworks. * Clean up silly error in finding the Kerberos Service Principal when using Kerberos authenticaion * Print useful error message when a connection unexpectedly closes, as this is usually authentication related... This commit was SVN r11923.	2006-10-01 22:43:17 +00:00
Brian Barrett	9807a38458	Always initialize the base output stream, but only set verbose if requested. Otherwise, the PLS components have pay more attention to debugging streams than the rest of the OMPI source code This commit was SVN r11922.	2006-10-01 22:37:30 +00:00
Ralph Castain	33a46cff40	Grrrr....ok, this time actually put a value in the silly command buffer! This commit was SVN r11901.	2006-09-29 21:44:11 +00:00
Ralph Castain	0625ead54a	Fix comm_spawn communications This commit was SVN r11900.	2006-09-29 21:10:48 +00:00
Ralph Castain	0411f9772e	Begin instrumenting for scalability tests. I have added a new MCA param (hey, you can't have too many!) called OMPI_MCA_orte_timing. If set to anything other than zero, the system will report out critical timing loops. At the moment, this includes three measurements: 1. Time spent going through the RDS->RAS->RMAPS, setting up triggers, etc. prior to calling the actual PLS launch function. This is reported out as time to setup job. 2. Time spent in MPI_Init from start of that function (well, right after opal_init) to the place where we send all of our info the registry. Reported out as time from start to exec_compound_cmd 3. Time actually spent executing the compound cmd. Reported out as time to exec_compound_cmd. A few additional timing points will be added shortly. These may eventually be removed or (better) setup with a conditional compile flag. This commit was SVN r11892.	2006-09-29 13:19:44 +00:00
Ralph Castain	db6a93fa63	Fix a couple of reported issues: 1. PLS finalize was not being called. Now ensure that happens during orte_finalize. 2. Errmgr proxies were sending their messages to the wrong tag - typical cut/paste error. This commit was SVN r11891.	2006-09-29 12:45:50 +00:00
Tim Prins	1b35e7adff	cleanup This commit was SVN r11863.	2006-09-28 13:28:48 +00:00

... 3 4 5 6 7 ...

1010 Коммитов