Commits · 886df85b9a561cb7c3658f0e4aafdaaf1f5fbdd9 · Manuel G. Marciani / ces_slurm_simulator

18 Mar, 2016 2 commits

Fix typo. · 886df85b
Tim Wickberg authored Mar 18, 2016

886df85b

Fix for srun abort on SIGSTOP+SIGCONT · 1ed38f26

Morris Jette authored Mar 18, 2016

Avoid possibly aborting srun that gets simultaneous SIGSTOP+SIGCONT while
    creating the job step. The result is that the signal hanlder gets a
    argument (the signal received) of zero.

Here's a log, window 1:
$ srun hostname
srun: Job step creation temporarily disabled, retrying
srun: I Got signal 18
srun: I Got signal 18
srun: I Got signal 18
srun: I Got signal 18
srun: I Got signal 18
srun: I Got signal 18
srun: I Got signal 18
srun: I Got signal 18
srun: I Got signal 18
srun: I Got signal 18
srun: I Got signal 18
srun: I Got signal 18
srun: I Got signal 0
srun: Cancelled pending job step

Window 2:
$  kill -STOP 18696 ; kill -CONT 18696
$  kill -STOP 18696 ; kill -CONT 18696
$  kill -STOP 18696 ; kill -CONT 18696
....

bug 2494

1ed38f26

17 Mar, 2016 2 commits

Merge branch 'slurm-14.11' into slurm-15.08 · dd2324a7
Tim Wickberg authored Mar 17, 2016
```
Update NEWS as well.
```
dd2324a7

Prevent uid update from corrupting assoc_hash table. · 60b58b70

Tim Wickberg authored Mar 17, 2016

The uid is used as part of the hash function, must remove old reference
and recalculate if it may change, otherwise _delete_assoc_hash
will not find it again when the association is removed, causing
slurmctld to segfault.

Bug 2560.

60b58b70

16 Mar, 2016 6 commits

Update gang scheduling data structures when job changes in size · 701917cc

Morris Jette authored Mar 16, 2016

Previous gang scheduling logic maintained information about resources
  originally allocated to the job and made scheduling decisions on
  that basis.
bug 2494

701917cc

Add signal number to error message · 83400184
Morris Jette authored Mar 16, 2016
```
This will improve ability to diagnose problems if the srun is
killed by a signal.
```
83400184

gang scheduling for with manually job suspend/resume · 344d2eab

Morris Jette authored Mar 16, 2016

Update gang scheduling table when job manually suspended or resumed. Prior
    logic could mess up job suspend/resume sequencing.
bug 2494

344d2eab

Fix issue when adding a new TRES to AccountingStorageTRES for the first · 6c436e34

Danny Auble authored Mar 16, 2016

time.

https://bugs.schedmd.com/show_bug.cgi?id=2547

The code just wasn't fully baked before and was probably written before
a lot of the other supporting code was done i.e
assoc_mgr_set_assoc|qos_tres_cnt were done specifically for this kind of
thing.  Many of the usage structures weren't realloced either as well as
the tres_cnt local to each qos and assoc wasn't updated.  So all in all
pretty bad code - bad Danny.  This makes sure all this sets up and no
memory corruption happens.

6c436e34

Send burst buffer teardown immediately · d85cdcc7

Morris Jette authored Mar 16, 2016

Generate burst buffer use completion email immediately afer teardown
    completes rather than at job purge time (likely minutes later).
bug 2539

d85cdcc7

Modify burst buffer stage out message · fae4c3d3

Morris Jette authored Mar 16, 2016

Change burst buffer use completion message from
"SLURM Job_id=1360353 Name=tmp Staged Out, StageOut time 00:01:47" to
"SLURM Job_id=1360353 Name=tmp StageOut/Teardown time 00:01:47"

fae4c3d3

15 Mar, 2016 5 commits
- acct_gather_energy/ipmi - add threshold for message logging · 18608974
  Alejandro Sanchez authored Mar 15, 2016
  
  18608974
- Document how to create slurmstepd core file · 4305fb7c
  Morris Jette authored Mar 15, 2016
  
  4305fb7c
- Clarify language for NoReserve flag. · 3c98f608
  Tim Wickberg authored Mar 15, 2016
```
Bug 2548. No functional change, documentation only.
```
  3c98f608
- Continue 5708037d with checks for tres_pos > 0. · 7274ef9f
  Tim Wickberg authored Mar 15, 2016
```
Otherwise "not found" value of -1 for tres_pos would cause
out-of-bounds memory access.
```
  7274ef9f
- Check that bb_state.tres_pos is set correctly to avoid overwriting CPU TRES. · 5708037d
  Tim Wickberg authored Mar 15, 2016
```
Bug 2543.
```
  5708037d
14 Mar, 2016 3 commits
- Change NoInAddrAnyCtld to NoCtldInAddrAny so as to not have it also · f5b5e605
  Danny Auble authored Mar 14, 2016
```
resolve NoInAddrAny when doing a strstr.  Continuation of commit 775c46de.
```
  f5b5e605
- Add option for TopologyParam=NoInAddrAnyCtld to make the slurmctld listen · 775c46de
  Danny Auble authored Mar 14, 2016
```
on only one port like TopologyParam=NoInAddrAny does for everything else.
```
  775c46de
- FreeBSD - set_oom_adj is Linux-specific, stub out to avoid errors. · b3f2359f
  Tim Wickberg authored Mar 14, 2016
```
There's no /proc on *BSD, and BSD handles OOM in a completely different way.
```
  b3f2359f
11 Mar, 2016 2 commits
- Merge branch 'slurm-14.11' into slurm-15.08 · a912fb3b
  Tim Wickberg authored Mar 11, 2016
  
  a912fb3b
- Fix job array step function printout. · 03d29e24
  Tim Wickberg authored Mar 11, 2016
```
Return [0-100:2] formatting, rather than [0,2,4,6,8,...] when using
a step function.

Was inadvertantly broken in 14.11 with commit 5ffdca92.

Bug 2535.
```
  03d29e24
10 Mar, 2016 2 commits

Add NEWS for commit 3bb2e602 · a0be0dc5
Morris Jette authored Mar 10, 2016

a0be0dc5

Cray Datawarp job requeue bug fix · 3bb2e602

Morris Jette authored Mar 10, 2016

burst_buffer/cray plugin: Prevent a requeued job from being restarted while
    file stage-out is still in progress. Previous logic could restart the job
    and not perform a new stage-in.
bug 2584, comment #45

3bb2e602

09 Mar, 2016 2 commits

cray job requeue bug · fec5e03b

Morris Jette authored Mar 09, 2016

Fix Cray NHC spawning on job requeue. Previous logic would leave nodes
allocated to a requeued job as non-usable on job termination.

Specifically, each job has a "cleaning/cleaned" flag. Once a job
terminates, the cleaning flag is set, then after the job node health
check completes, the value gets set to cleaned. If the job is requeued,
on its second (or subsequent) termination, the select/cray plugin
is called to launch the NHC. The plugin sees the "cleaned" flag
already set, it then logs:
error: select_p_job_fini: Cleaned flag already set for job 1283858, this should never happen
and returns, never launching the NHC. Since the termination of the
job NHC triggers releasing job resources (CPUs, memory, and GRES),
those resources are never released for use by other jobs.

Bug 2384

fec5e03b

Correctly parse nids in slurmconfgen_smw.py · 88ccc111

David Gloe authored Mar 09, 2016

An error in slurmconfgen_smw.py caused it to parse the nic as the nid.
On some systems those values differ, causing the generated slurm.conf file to
be incorrect.

Bug 2532.

88ccc111

08 Mar, 2016 5 commits
- Remove unneeded check introduced in 897c4b27 · ba7dfc75
  Tim Wickberg authored Mar 08, 2016
```
_set_collectors() already has a run_in_daemon("slurmd") that
precludes this from being an issue.
```
  ba7dfc75
- Fix route/topology plugin to prevent segfault in sbcast. · 897c4b27
  Bill Brophy authored Mar 08, 2016
```
route_p_split_hostlist was not thread-safe, and would cause
one of several segfaults depending on where in the initialization
code each thread was.

Bug 2495.
```
  897c4b27
- Fix displayed value for RoutePlugin. · 14c51e65
  Tim Wickberg authored Mar 08, 2016
```
Was incorrectly displaying "(null)" even when loaded successfully.
```
  14c51e65
- Handle function error for Coverity · e6b7b2c2
  Morris Jette authored Mar 08, 2016
  
  e6b7b2c2
- Fix link to kernel cgroup documentation. · 4c3fd194
  Janne Blomqvist authored Mar 08, 2016
  
  4c3fd194
07 Mar, 2016 1 commit

add additional tuning notes for mysql/mariadb · 49dc5d8d

Tim Wickberg authored Mar 07, 2016

In particular, it seems that MariaDB has changed the default for
innodb_lock_wait_timeout has been lowered which can cause issues
for the various rollup processes on systems with high job counts.

49dc5d8d

05 Mar, 2016 2 commits
- Continuation to commit b294f81b to do the right thing for jobs. · 35f7a262
  Danny Auble authored Mar 04, 2016
  
  35f7a262
- Fixed double read lock on getting job's gres/tres. · b23a57cf
  Danny Auble authored Mar 04, 2016
  
  b23a57cf
04 Mar, 2016 2 commits
- Continuation of commit 7f0bdc84 · 55a678dd
  Danny Auble authored Mar 04, 2016
```
Step GRES value changed from type "int" to "int64_t" to support larger
values.

Signed-off-by: Danny Auble <da@schedmd.com>
```
  55a678dd
- Fix issue where steps weren't always getting the gres/tres involved. · b294f81b
  Danny Auble authored Mar 04, 2016
  
  b294f81b
03 Mar, 2016 4 commits

Fix issue with sbcast not doing a correct fanout. · 72f13426
Danny Auble authored Mar 03, 2016

72f13426
Fix getting reservations to database when database is down. · 5c43d754
Brian Christiansen authored Mar 03, 2016
```
Bug 2507
```
5c43d754

Increase step GRES variable size · 7f0bdc84

Morris Jette authored Mar 03, 2016

Step GRES value changed from type "int" to "int64_t" to support larger
values. Previous logic could fail in step allocation values over 32-bits.
Other GRES values are 64-bit.

7f0bdc84

Force close on exec on first 256 file descriptors when launching a · f502f1e5

Danny Auble authored Mar 02, 2016

slurmstepd to close potential open ones.

It was pointed out the slurmd using acct_gather_energy/ipmi links to
freeipmi which could possibly open /dev/ipmi0 without the close on exec
flag set as root while launching a step leaving it open in the users app.

What this does is sets the flag on the first 256 to mitigate the concern.

Reported by Maksym Planeta.

Bug 2506

f502f1e5

02 Mar, 2016 2 commits

Backfill scheduler to validate correct job partition · efd9d35e

Gary B Skouson authored Mar 02, 2016

Previous logic tested whatever the job's partition pointer indicated
rather than the partition we are trying to run the job in. This bug
was introduced in Slurm version 15.08.5, Nov 16, 2015, commit
94f0e948
bug 2499

efd9d35e

Move definition to only place used to avoid confusion, continuation of · f257976a
Danny Auble authored Mar 02, 2016
```
patch 2d5066e7
```
f257976a