openmpi

Автор	SHA1	Сообщение	Дата
Ralph Castain	2110fb7f95	Add some debug This commit was SVN r27257.	2012-09-07 04:06:37 +00:00
Ralph Castain	78ccb097f0	Fix vm setup in unmanaged environments - needs to construct a node list in the same way we now do for mapping This commit was SVN r27256.	2012-09-07 01:53:19 +00:00
Ralph Castain	876d78f36a	If JAVA_HOME is present on a Linux system, use it to find Java support This commit was SVN r27255.	2012-09-07 00:41:14 +00:00
Ralph Castain	36acbe4ca6	Multiple apps might want the same files, so instead of using the app_idx to determine who gets what, use the actual file names as they are sent anyway This commit was SVN r27254.	2012-09-06 22:02:05 +00:00
Ralph Castain	e9e52fc78f	Gain some efficiency in the staged mapper - if soft locations are in use and get_nodes returns busy, then no need to continue cycling thru the remaining apps as all nodes are occupied This commit was SVN r27253.	2012-09-06 22:01:18 +00:00
Ralph Castain	67f34c3be6	Record the bind_level recvd by the daemon for each job so it can be correctly sent to the procs. Add test in get_relative_locality to avoid descending into an infinite loop if the level is NODE (==0). This commit was SVN r27252.	2012-09-06 20:50:07 +00:00
Jeff Squyres	dd254cc202	OMPI_HAVE_IBV_LINK_LAYER does not exist. Instead, check defined(HAVE_IBV_LINK_LAYER_ETHERNET). This commit was SVN r27251.	2012-09-06 18:25:36 +00:00
Jeff Squyres	aca005ccc5	Add bullet about MPI_CART_SUB This commit was SVN r27249.	2012-09-06 14:24:02 +00:00
Jeff Squyres	0d2962ebf0	Fixes trac:3294: space for the periods has already been allocated by ompi_comm_split(), and the entire set of periods from the old communicator have already been copied to the new communicator. But up here in mca_topo_base_cart_sub(), we need to subset the periods that are actually stored on the new communicator according to remain_dims (just like we did for the set of dimensions). This commit renames a few variables to be a little less misleading, and then adds a loop to copy over the periods information. I could have added this into the first loop (that subset-copies the dimensions), but this code is already confusing enough and this is not a performance-critical section: so I made it a new loop. Note that all the topo code will be revamped a bit when the new MPI-2.2 topo stuff (currently off in a mercurial branch) finally makes it back to the SVN trunk. But that new stuff will only get to v1.7 -- this commit will need to be CMR'ed to v1.6.x. cmr:v1.7 cmr:v1.6.2 This commit was SVN r27248. The following Trac tickets were found above: Ticket 3294 --> https://svn.open-mpi.org/trac/ompi/ticket/3294	2012-09-06 14:16:29 +00:00
Vishwanath Venkatesan	b75d877a3f	Removing .ompi_ignore for the lustre component. This commit was SVN r27247.	2012-09-05 22:20:18 +00:00
Vishwanath Venkatesan	640aca6654	Modifying the file view generation to remove the merging of offset-length pair. Its no longer needed as the default file view makes sure the chunks are large enough. This commit was SVN r27246.	2012-09-05 21:00:47 +00:00
Ralph Castain	efa50346c8	Error out if we are filtering a hostfile and encounter a node that is not in the resource-managed allocation, giving an error message identifying the file and the node. Don't filter managed allocations thru a default hostfile as this can lead to "hidden" errors. Don't use dash-host info on managed allocations if we using soft locations This commit was SVN r27245.	2012-09-05 19:42:00 +00:00
Brian Barrett	fa4c2af9ed	THe Portals 4 reference implementation will sometimes return a NI_FLOWCTL for both a send and an ack. I'm not sure whether this violates the spec, so work around until we decide... This commit was SVN r27244.	2012-09-05 19:36:19 +00:00
Jeff Squyres	38440369a7	Add note about Absoft compiler and the mpi_f08 module. This commit was SVN r27243.	2012-09-05 18:45:38 +00:00
Ralph Castain	d772e0fc3d	Add an option to treat dash-host specifications as "requested, but not required". So-called "soft" location requests can allow an application to execute even if the ideal allocation isn't available. This commit was SVN r27242.	2012-09-05 18:42:09 +00:00
Ralph Castain	6d29cecce1	Fix the help message warning of multiple prefixes so it correctly prints out the info, and fix a typo. cmr:v1.7 This commit was SVN r27241.	2012-09-05 16:28:36 +00:00
Ralph Castain	64ccf789f2	Ensure the final output is printed cmr:v1.7 This commit was SVN r27240.	2012-09-05 16:25:15 +00:00
Ralph Castain	fde83a44ab	This confusion has been around for awhile, caused by a long-ago decision to track slots allocated to a specific job as opposed to allocated to the overall mpirun instance. We eliminated that quite a while ago, but never consolidated the "slots_alloc" and "slots" fields in orte_node_t. As a result, confusion has grown in the code base as to which field to look at and/or update. So (finally) consolidate these two fields into one "slots" field. Add a field in orte_job_t to indicate when all the procs for a job will be launched together, so that staged operations can know when MPI operations are allowed. This commit was SVN r27239.	2012-09-05 01:30:39 +00:00
Ralph Castain	bae5dab916	If (and only if) a user requests, set the default number of slots on any node to the number of objects of the specified type. This only takes effect in an unmanaged environment - i.e., if an external resource manager assigns us a number of slots, then that is what we use. However, if we are using a hostfile, then the user may or may not have given us a value for the number of slots on each node. For those nodes (and only those nodes) where the user does not specify a slot count, we will set the number of slots according to their direction: either to the number of cores, numas, sockets, or hwthreads. Otherwise, the slot count is set to 1. Note that the default behavior remains unchanged: in the absence of any value for #slots, and in the absence of any directive to set #slots, we will set #slots=1. This commit was SVN r27236.	2012-09-04 20:58:26 +00:00
Ralph Castain	ee6c7702d2	Ensure the cma.h file is included in the tarball This commit was SVN r27235.	2012-09-04 19:34:09 +00:00
Ralph Castain	18d2f75b56	Ensure we don't re-link files when staging execution as we may be executing more members of the same app. Allow the user to ask that directory trees be "flattened" so that all files appear in the proc's session directory itself. This commit was SVN r27232.	2012-09-04 17:52:12 +00:00
Ralph Castain	86a16edd5a	Move the debug output to the right place This commit was SVN r27231.	2012-09-04 17:22:17 +00:00
Ralph Castain	11de735e8a	Complete the revamp of hostfile support in non-managed environments. Working at the app level, ensure that we utilize only those nodes specified for that app, but fall back to the default hostfile (if available) for those with no specification, further falling back to the local host if the default hostfile is not present or is empty. This commit was SVN r27230.	2012-09-04 16:34:05 +00:00
Jeff Squyres	7fbaaff94e	Always define OMPI_HAVE_RDMAOE; don't define it conditionally. This commit was SVN r27229.	2012-09-04 15:49:09 +00:00
Jeff Squyres	9feb8d8879	Oops; the error paths were not correct on the initial commit. Fixed. This commit was SVN r27228.	2012-09-04 15:48:44 +00:00
Shiqing Fan	0326e88c51	As opal_hwloc_topo_data_t has to create a class instance in orte, its definition has to be exported. Otherwise, there will be unresolved variable error on Windows. This commit was SVN r27227.	2012-09-04 13:52:29 +00:00
Jeff Squyres	b23a6b8eda	Shiqing removed this file in r27217 (but neglected to remove it from the Makefile.am). This commit was SVN r27226. The following SVN revision numbers were found above: r27217 --> open-mpi/ompi@ddbd542732	2012-09-04 13:06:39 +00:00
Ralph Castain	42d58f17bf	Provide a more user-expected way of handling hostfile and dash-host allocations in unmanaged environments. First look for an RM-managed allocation. If nothing is found, or no active module is alive, proceed to look for hostfile and dash-host assignments. These are provided on a per-app basis, so we have to cycle across the apps using the following algorithm: 1. if a hostfile is given, then add the nodes found in that hostfile to our list - i.e., the resulting allocation contains the UNION of all nodes specified in hostfiles from across all apps. 2. any app that has no hostfile but has a dash-host, will have those nodes added to the list 3. any app that fails to have a hostfile or a dash-host will be given the default hostfile, if we have it Each app will subsequently be filtered using their hostfile and/or dash-host data to ensure that the app only has access to the hosts it specified Note that any relative node syntax found in the hostfiles or dash-host data will generate an error in this scenario, so only non-relative syntax can be present This commit was SVN r27223.	2012-09-04 11:50:52 +00:00
Ralph Castain	3894179e2f	Add missing file This commit was SVN r27222.	2012-09-04 01:16:58 +00:00
Ralph Castain	fa6a18f05a	If the orted HNP is spawned by a singleton, then it needs to harvest all the MCA params from its environment to ensure they are passed on to any subsequently spawned daemons. Otherwise, the singleton could be directed to select options that other apps miss. This commit was SVN r27221.	2012-09-04 01:10:26 +00:00
Ralph Castain	b5e26a90ea	Singletons should insist that the spawned orted HNP use the novm state machine so that any subsequent comm_spawn occurs only on required nodes This commit was SVN r27220.	2012-09-04 01:09:01 +00:00
Ralph Castain	8de1291303	Set ignores This commit was SVN r27219.	2012-09-04 00:54:18 +00:00
Shiqing Fan	3bdb207cae	Revert the example VS project file. Update the CMakeLists.txt for the example, make sure it won't modify the original project file. This commit was SVN r27218.	2012-09-03 10:42:52 +00:00
Shiqing Fan	ddbd542732	Remove one .windows file. Add a macro definition for isblank function. This commit was SVN r27217.	2012-09-03 09:51:44 +00:00
Aleksey Senin	33ae1fe6c7	Fix untitialized return code in ompi_mtl_mxm_add_procs function. This commit was SVN r27216.	2012-09-02 13:17:49 +00:00
Yevgeny Kliteynik	3fe239702a	Fixed compilation error Thanks to Alex Margolin for the fix This commit was SVN r27215.	2012-09-02 08:26:30 +00:00
Ralph Castain	66c3f5d18d	When getting target nodes for mapping, there is a difference between not finding any nodes that match the required constraints (either in hostfile or dash-host filtering) and finding at least one such node, but all its slots are busy. Make the return code reflect this difference so the caller can take appropriate action. This commit was SVN r27213.	2012-09-01 10:30:40 +00:00
Jeff Squyres	341ce2f9a4	Per some discussions between LANL, Cisco, ORNAL, and Mellanox, move some new common OpenFabrics functionality to ompi/mca/common/verbs. Also move everything that was in ompi/mca/common/ofautils under ompi/mca/common/verbs. * Move ofautils -> verbs * Add new functionality in ompi/mca/common/verbs (see doxygen * comments in ompi/mca/common/verbs/common_verbs.h for details): * ompi_common_verbs_find_ibv_ports() * ompi_common_verbs_port_bw() * ompi_common_verbs_mtu() * '''If you're writing verbs-based code, you should be using this common functionality''' * Adapt openib BTL to use some trivial common functionality in common/verbs * Don't use "#ifdef OMPI_HAVE_RDMAOE",use "#if defined(HAVE_IBV_LINK_LAYER_ETHERNET)" * Update the following to include/link against common/verbs * bcol/iboffload * sbgp/ibnet * btl/openib This commit was SVN r27212.	2012-09-01 01:42:37 +00:00
Shiqing Fan	b27862e5c7	make sure the static build on windows can link the required libraries. This commit was SVN r27211.	2012-08-31 23:27:11 +00:00
Ralph Castain	95019cc310	Fix a few places where we weren't completely identifying hostfile-based operations against "localhost" entries. Tell the mapper base to be silent when we don't want errors announced because nodes aren't available for mapping (something it is okay if they are fully used). Fix an infinite loop in the file prepositioning code. This commit was SVN r27210.	2012-08-31 21:28:49 +00:00
Pavel Shamis	888b04ab36	Fixing gcc 4.7.1 warning in ptpcoll bcol. Refs trac:3243. This commit was SVN r27209. The following Trac tickets were found above: Ticket 3243 --> https://svn.open-mpi.org/trac/ompi/ticket/3243	2012-08-31 21:16:58 +00:00
Jeff Squyres	da00d281e6	Oops -- we want the priority to be low, not high (for now). :-) This commit was SVN r27208.	2012-08-31 21:08:35 +00:00
Jeff Squyres	36dc0d40a6	* Fix a few warnings in ompi_rb_tree * Add the get_key function to the opal_tree test This commit was SVN r27207.	2012-08-31 20:43:58 +00:00
Jeff Squyres	fcc1c7e33c	= Overview = First revision of the Locatation Aware Mapping Algorithm (LAMA) RMAPS component. This component is used to effect many different types of regular of process/processor affinity patterns. Although quite flexible in the patterns that it provides, it is ''not'' a fully-arbitrary, rankfile-like solution for process/processor affinity. Inspiried by !BlueGene-like network specifications, LAMA has a core algorithm that is quite good at specifying regular patterns in multiple "dimensions" (where "dimensions" are expressed in terms of different hardware elements: processor hardware threads, cores, sockets, ...etc.). The LAMA core algorithm is described here: http://www.open-mpi.org/papers/cluster-2011-lama/ = LAMA Usage Levels = LAMA allows specifying affinity multiple different ways: 1. None: Speciying no affinity options to mpirun results in exactly the same behavior as today: no affinity is used. 1. Simple: Using the mpirun options "--bind-to <WIDTH>" and "--map-to <LEVEL>" to indicate how "wide" each process should be bound (i.e., bind to a processor core, or to a processor socket, etc.) and how to lay out the processes (i.e., round robin by cores, sockets, etc.). 1. Expert: Using four new MCA parameters to effect process mapping and binding to processors. These options are a bit complex, and are not for the faint at heart, but offer a high degree of (regular pattern) flexibility (each of these are described more fully below): * rmaps_lama_map: a sequence of characters describing how to lay out processes * rmaps_lama_bind: a sequence of characters describing the resources to bind to each process * rmaps_lama_mppr: a sequence of characters describing the maximum number of processes to allow per resource (i.e., a specific definition of "oversubscription") * rmaps_lama_ordering: once all processes are in place, how to order the ranks in MPI_COMM_WORLD We anticipate that most users will utilize the "None" and "Simple" levels of affinity, and they continue to work just as they do with the v1.6 series and SVN trunk. The Expert level was designed for two purposes: 1. To provide a precise definition for the "Simple" level (i.e., every --bind-to/--map-by option in the "Simple" level has a corresponding precise specification in the "Expert" level) 1. As modern computing platforms become more complex, we simply cannot predict what application developers will need in terms of processor affinity. LAMA is an attempt to provide a highly flexible mechanism that allows applications to utilize a variety of complex, unique affinity patterns beyond the common "bind to core" and "bind to socket" patterns. = LAMA Simple Level = The "Simple" level is pretty much the same as what Open MPI has offered for years. It supports the same --bind-to and --map-by options that Open MPI has supported for a while, but expands their scope a bit. Specifically, the following options are available for both --bind-to and --map-by: * slot * hwthread * core * l1cache * l2cache * l3cache * socket * numa * board * node = LAMA Expert Level = The "Expert" level requires some explanation. I'll repeat my disclaimer here: the LAMA Expert level is not for the meek. It is flexible, but complex. '''Most users won't need the Expert level.''' LAMA works in three phases: mapping, binding, and ordering. Each is described below. == Expert: Mapping == Processes are paired with sets of resources. For example, each process may be paired with a single processor core. Or each process may be paired with an entire processor socket. LAMA performs this mapping, obeying the Max Processes Per Resource ("MPPR", pronounced "mipper") limits. More on MPPR, below. Mapping can be performed across multiple hardware levels: * h: Hardware thread * c: Processor core * s: Processor socket * L1: L1 cache * L2: L2 cache * L3: L3 cache * N: NUMA node * b: Processor board * n: Server node If the act of mapping is that of pairing MPI processes to the resources that have been allocated to a job, one can easily imagine looping through all the resources and assigning processes to them. But to effect different process process layout patterns across those resources, one may want to loop over those resources ''in a different order.'' That is, if the above-mentioned nine hardware resources (hardware thread, processor core, etc.) can be thought of as an nine-dimensional space, you can imagine nine nested loops to traverse all of them. And you can imagine that changing the order of nesting would change the traversal pattern. LAMA accepts a sequence of tokens representing the above-mentioned nine hardware resources to specify the order of looping when mapping resources to processes. For example, consider a "simple" traversal: csL1L2L3Nbnh. Reading that sequence of letters from left-to-right, it specifies mapping by processor core, processor socket, L1 cache, L2 cache, L3 cache, NUMA node, processor board, server node, and finally hardware thread. Wait... what? That string specifies resources from "smallest" to "largest" -- with the exception of hardware threads. Why are they tacked on to the end? In short, this string of letters means "map by round robin by core" -- (indeed, it exactly corresponds to the Simple level "--map-by core"). Specifically, LAMA traverses the string from left-to-right and maps processes to all the resources indicated by that token (e.g., "c" for processor core). When there are no more resources indicated by that token, it goes on to the next token. Hence, in this case, LAMA will map the first process to the first core, then it will map the second process to the second core, and so on. Once all the cores are exhausted, LAMA effectively ignores all the other letters until "h" (because all the other resources are made up of cores; when cores are exhausted, those resources are exhausted, too). If there are still more processes to be mapped, LAMA will then traverse all the hyperthreads -- meaning that the next process will be mapped to the second hyperthread on the first core. And the next process will be mapped to the second hyperthread on the second core. And so on. Keep in mind that the cores involved may span many server nodes; we're not just talking about the cores (etc.) in a single machine. As another example, the sequence "sL1L2L3Nbnch" is exactly equivalent to "--map-by socket" (i.e., LAMA maps the first process to the first socket, the second process to the second socket, and so on). The sequence of letter can be combined in many, many different ways to produce many different regular mapping patterns. === Max Processes Per Resource (MPPR) === The MPPR is an expression that precisely defines the maximum number of processes that can be mapped to any single resource. In effect, it defines the concept of "oversubscription." Specifically, traditional HPC wisdom is that "oversubscription" is when there is more than one MPI process per processor core. This conventional defintion is expressed in a MPPR string of "1:c" (one process per core). But what if your MPI processes are multi-threaded, and they need multiple processes per core? You'd need a different description of "oversubscription" in this case. Perhaps you want to have one MPI process per socket. This would be expressed in a MPPR string of "1:s". The general form of an individual MPPR specification is an integer follow by a colon, followed by any of the tokens from mapping can be used in the MPPR specification. For example "1:c" is pronounced "one process per core." Multiple MPPR specifications can be strung together into a comma-delimited list, too. All of these MPPR values and then taken into account when mapping. Here's some examples: * 1:c -- allow, at most, one process per processor core (i.e., don't schedule by hyperthread) * 1:s -- allow, at most, one process per processor socket (e.g., that process may be multithreaded, or wants exclusive use of the socket's caches) * 1:s,2:n -- only allow one process per processor socket, but, at most, two processes per server node (e.g., if the two MPI processes will consume all the RAM on the server node, even if there are more processor cores available) If mapping all processes to resources would exceed a MPPR limit, this job is ruled to be oversubscribed. If --oversubscribe was specified on the mpirun command line, the job continues. Otherwise, LAMA will abort the job. Additionally, if --oversubscribe is specified, LAMA will endlessly cycle through the mapping token string untill all processes have been mapped. == Expert: Binding == Once processes have been paired with resources during the Mapping stage, they are optionally bound to a (potentially different) set of resources. For example, processes may be mapped round robin by processor socket, but bound to an individual processor core. To be clear: if binding is not used, then mapping is effectively reduced to "counting how many processes end up on each server node." Without binding, there's no enforcement that a process will stay where LAMA thinks it was placed. With binding, however, processes are bound to a set of hardware threads. The number of threads to which the process is bound is sometimes referred to as the "binding width". For example, if a process is bound to all the hardware threads in a processor socket, its "width" is the processor socket. (note that we specifically do not say that the hardware threads are sequential, even if they are all within a single resource such as a processor core or socket. BIOS ordering of hardware threads can be wonky; so we only refer to "sets of hardware threads") Bindings are expressed as an integer and a token from the mapping string. For example "1s" means "bind each process to one processor socket" (there is no ":" in the binding string because the ":" is pronounced as "per" when reading the MPPR string). Note that it only makes sense to bind processes to a single resource specification (unlike the MPPR specification, where multiple limits can be specified). == Expert: Ordering == Finally, processes are assigned a rank in MPI_COMM_WORLD. LAMA currently offers two ordering modes: sequential or natural: * Sequential: if you laid out all the hardware resources in a single line, and then overlaid all the MPI processes on top of them, they are ordered from 0 to (N-1) from left-to-right. * Natural: the ordering of ranks follows the mapping ordering. For example, consider a server node with two processor sockets, each containing four cores. The command line "mpirun -np 8 --bind-to core --map-by socket --order n a.out" would result in MCW ranks that look like this: [0 2 4 6] [1 3 5 7]. = Execution = At this point, the job is fully mapped, optionally bound, and its ranks in MPI_COMM_WORLD are ordered. It now starts its execution. = Final Notes = Note that at this point, lama is not the default mapper. It must be activiated with "--mca rmaps lama". We'll continue to do further testing and comparitive analysis with the current set of ORTE mappers. Also, note that the LAMA algorithm can handle heterogeneity between hardware resources (e.g., an MPI job spanning server nodes with differing numbers of processor sockets). For lack of a longer explanation (this commit message already long enough!), LAMA considers each server node individually during mapping and binding. See the LAMA paper for more details: http://www.open-mpi.org/papers/cluster-2011-lama/ This commit was SVN r27206.	2012-08-31 19:57:53 +00:00
Jeff Squyres	287e47a04d	Fixes, improvements, and enhancements to the opal_tree class (used by the LAMA RMAPS component, to be committed shortly). This commit was SVN r27204.	2012-08-31 16:35:49 +00:00
Jeff Squyres	ca5bd85364	Add some help messages to the RAS simulator for PEBKAC issues. This commit was SVN r27203.	2012-08-31 16:33:29 +00:00
Josh Hursey	f1b43f6375	Add UW-L to the License file This commit was SVN r27202.	2012-08-31 16:12:40 +00:00
Josh Hursey	bdcf1717dd	Fix an unchecked return code - Thanks Jeff S. for noticing it. This commit was SVN r27201.	2012-08-31 16:12:03 +00:00
Jeff Squyres	de512db1fd	Simpler scheme than r27195: if the gk commit file ends up being 0 bytes long, then abort the commit. This avoids asking an extra question in the most common case (where the GK doesn't edit the file at all). This commit was SVN r27198. The following SVN revision numbers were found above: r27195 --> open-mpi/ompi@70aa879ed3	2012-08-31 16:05:09 +00:00
Ralph Castain	6319014ab0	Sigh - get the end of the loop at the right place This commit was SVN r27197.	2012-08-31 15:54:11 +00:00

... 3 4 5 6 7 ...

17660 Коммитов