2006-06-06 01:24:42 +04:00
|
|
|
# -*- text -*-
|
|
|
|
#
|
|
|
|
# Copyright (c) 2004-2005 The Trustees of Indiana University and Indiana
|
|
|
|
# University Research and Technology
|
|
|
|
# Corporation. All rights reserved.
|
|
|
|
# Copyright (c) 2004-2005 The University of Tennessee and The University
|
|
|
|
# of Tennessee Research Foundation. All rights
|
|
|
|
# reserved.
|
|
|
|
# Copyright (c) 2004-2005 High Performance Computing Center Stuttgart,
|
|
|
|
# University of Stuttgart. All rights reserved.
|
|
|
|
# Copyright (c) 2004-2006 The Regents of the University of California.
|
|
|
|
# All rights reserved.
|
|
|
|
# $COPYRIGHT$
|
|
|
|
#
|
|
|
|
# Additional copyrights may follow
|
|
|
|
#
|
|
|
|
# $HEADER$
|
|
|
|
#
|
|
|
|
# This is the US/English general help file for Open MPI.
|
|
|
|
#
|
|
|
|
[btl_openib:retry-exceeded]
|
2006-06-20 15:23:38 +04:00
|
|
|
The InfiniBand retry count between two MPI processes has been
|
|
|
|
exceeded. "Retry count" is defined in the InfiniBand spec 1.2
|
|
|
|
(section 12.7.38):
|
2006-06-06 01:24:42 +04:00
|
|
|
|
2006-06-20 15:23:38 +04:00
|
|
|
The total number of times that the sender wishes the receiver to
|
|
|
|
retry timeout, packet sequence, etc. errors before posting a
|
|
|
|
completion error.
|
2006-06-06 06:04:56 +04:00
|
|
|
|
2006-06-20 15:23:38 +04:00
|
|
|
This error typically means that there is something awry within the
|
|
|
|
InfiniBand fabric itself. You should note the hosts on which this
|
|
|
|
error has occurred; it has been observed that rebooting or removing a
|
|
|
|
particular host from the job can sometimes resolve this issue.
|
2006-06-06 06:04:56 +04:00
|
|
|
|
2006-06-20 15:23:38 +04:00
|
|
|
Two MCA parameters can be used to control Open MPI's behavior with
|
|
|
|
respect to the retry count:
|
2006-06-06 06:04:56 +04:00
|
|
|
|
2006-06-20 15:23:38 +04:00
|
|
|
* btl_openib_ib_retry_count - The number of times the sender will
|
|
|
|
attempt to retry (defaulted to 7, the maximum value).
|
|
|
|
|
|
|
|
* btl_openib_ib_timeout - The local ACK timeout parameter (defaulted
|
|
|
|
to 10). The actual timeout value used is calculated as:
|
|
|
|
|
|
|
|
4.096 microseconds * (2^btl_openib_ib_timeout)
|
|
|
|
|
|
|
|
See the InfiniBand spec 1.2 (section 12.7.34) for more details.
|