Newer
Older
Christopher J. Morrone
committed
* Changes in SLURM 1.0.11
=========================
-- Fix for slurmstepd hang when launching a task. (Needed to install
list library's atfork handlers).
-- Fix memory leak on AIX (and possibly other architectures) due to
missing pthread_attr_destroy() calls.
-- Fix rare task standard I/O setup bug. When the bug hit, stdin, stdout,
Christopher J. Morrone
committed
or stderr could be an invalid file descriptor.
-- General slurmstepd file descriptor cleanup.
-- Fix memory leak in job accounting logic (Andy Riebs, HP, memory_leak.patch).
Christopher J. Morrone
committed
* Changes in SLURM 1.0.10
=========================
-- Fix for job accounting logic submitted from Andy Riebs to handle issues
with suspending jobs and such. patch file named requeue.patch
-- Make select/cons_res interoperate with mpi/lam plugin for task counts.
-- Fix race condition where srun could seg-fault due to use of logging functions
within pthread after calling log_fini.
-- Code changes for clean build with gcc 2.96 (gcc_2_96.patch, Takao Hatazaki, HP).
-- Add CacheGroups configuration support in configurator.html (configurator.patch,
Takao Hatazaki, HP).
-- Fix bug preventing use of mpich-gm plugin (mpichgm.patch, Takao Hatazaki, HP).
* Changes in SLURM 1.0.9
========================
-- Fix job accounting logic to open new log file on slurmctld reconfig.
(Andy Riebs, slurm.hp.logfile.patch).
-- Fix bug which allows a user to run a batch script on a node not allocated
by the slurmctld.
-- Fix poe MP_HOSTFILE handling bug on AIX.
* Changes in SLURM 1.0.8
========================
-- Fix to communication between slurmd and slurmstepd to allow for partial
reads and writes on their communication pipes.
* Changes in SLURM 1.0.7
========================
-- Change in how AuthType=auth/dummy is handled for security testing.
-- Fix for bluegene systems to allow full system partitions to stay booted
when other jobs are submitted to the queue.
* Changes in SLURM 1.0.6
========================
-- Prevent slurmstepd from crashing when srun attaches to batch job.
* Changes in SLURM 1.0.5
========================
-- Restructure logic for scheduling BlueGene small block jobs. Added
"test_only" flag to select_p_job_test() in select plugin.
-- Correct squeue "NODELIST" output for BlueGene small block jobs.
-- Fix possible deadlock situations on BlueGene plugin on errors.
* Changes in SLURM 1.0.4
========================
-- Release job allocation if step creation fails (especially for BlueGene).
-- Fix bug select/bluegene warm start with changed bglblock layout.
-- Fix bug for queuing full-system BlueGene jobs.
* Changes in SLURM 1.0.3
========================
-- Fix bug that could refuse to queue batch jobs for BlueGene system.
-- Add BlueGene plugin mutex lock for reconfig.
-- Ignore BlueGene bgljobs in ERROR state (don't try to kill).
-- Fix job accounting for batch jobs (Andy Riebs, HP,
slurm.hp.jobacct_divby0a.patch).
-- Added proctrack/linuxproc.so to the main RPM.
-- Added mutex around bridge api file to avoid locking up the api.
-- BlueGene mod: Terminate slurm_prolog and slurm_epilog immediately if
SLURM_JOBID environment variable is invalid.
-- Federation driver: allow selection of a sepecific switch interface
(sni0, sni1, etc.) with -euidevice/MP_EUIDEVICE.
-- Return an error for "scontrol reconfig" if there is already one in
progress
* Changes in SLURM 1.0.2
========================
-- Correctly report DRAINED node state as type OTHER for "sinfo --summarize".
-- Fixes in sacct use of malloc (Andy Riebs, HP, sacct_malloc.patch).
-- Smap mods: eliminate screen flicker, fix window resize, report more clear
message if window too small (Dan Palermo, HP, patch.1.0.0.1.060126.smap).
-- Sacct mods for inconsistent records (race condition) and replace --debug
option with --verbose (Andy Riebs, HP, slurm.hp.sacct_exp_vvv.patch).
-- scancel of a job step will now send a job-step-completed message
to the controller after verifying that the step has completed on all nodes.
-- Fix task layout bug in srun.
-- Added times to node "Reason" field when set down for insufficient
resources or if not responding.
-- Validate operation with Elan switch and heterogeneous nodes.
* Changes in SLURM 1.0.1
========================
-- Assorted updates and clarifications in documentation.
-- Detect which munge installation to use 32/64 bit.
-- Fix sinfo filtering bug, especially "sinfo -R" output.
-- Fix node state change bug, resuming down or drained nodes.
-- Fix "scontrol show config" to display JobCredentialPrivateKey instead
of JobCredPrivateKey and JobCredentialPublicCertificate instead of
JobCredPublicKey. They now match the options in the slurm.conf.
-- Fix bug in job accounting for very long node list records (Andy Riebs,
HP, sacct_buf.patch).
-- BLUEGENE SPECIFIC - added load function to smap to load an already
exsistant bluegene.conf file.
-- Fix bug in sacct: If user requests specific job or job step ID,
only the last one with that ID will be reported. If multiple
nodes fail, the job has its state recorded as "JOB_TERMINATED...nf"
(Andy Riebs, HP, slurm.hp.sacct_dup.patch).
-- Fix some inconsistencies in sacct's help message (Andy Riebs, HP,
slurm.hp.sacct_help.patch).
-- Validate input to sacct command and allows embedded spaces in
arguments (Andy Riebs, HP, slurm.hp.sacct_validate.patch).
* Changes in SLURM 0.7.0-pre8
=============================
-- BGL specific -- bug fix for smap configure function down configuration
Christopher J. Morrone
committed
-- Add slurmd cache for group IDs (Takao Hatazaki, HP).
-- Fix bug in processing of "#SLURM" batch script option parsing.
* Changes in SLURM 0.7.0-pre7
=============================
-- Fix issue with NODE_STATE_COMPLETING, could start job on node before
epilog completed.
-- Added some infrastructure for job suspend/resume (scontrol, api, and
slurmctld stub).
-- Set job's num_procs to the actual processor count allocated to the job.
-- Fix bug in HAVE_FRONT_END support for cluster emulation.
* Changes in SLURM 0.7.0-pre6
=============================
-- Added support for task affinity for binding tasks to CPUs (Daniel
Palermo, HP).
-- Integrate task affinity support with configuration, add validation
test.
* Changes in SLURM 0.7.0-pre5
=============================
-- Enhanced performance and debugging for slurmctld reconfiguration.
-- Add "scontrol update Jobid=# Nice=#" support.
-- Basic slurmctld and tool functionality validated to 16k nodes.
-- squeue and smap now display correct info for jobs in bluegene enviornment.
-- Fix setting of SLURM_NODELIST for batch jobs.
-- Add SubmitTime to job information available for display.
-- API function slurm_confirm_allocation() has been marked OBSOLETE
and will go away in some future version of SLURM. Use
-- New API calls slurm_signal_job and slurm_signal_job_step to send
signals directly to the slurmds without triggering the shutdown sequence.
-- remove "uid" from old_job_alloc_msg_t, no longer needed.
-- Several bug fixes in maui scheduler plugin from Dave Jackon
(Cluster Resources).
* Changes in SLURM 0.7.0-pre4
=============================
-- Remove BNR libary functions and add those for PMI (KVS and basic
MPI-1 functions only for now)
-- Added Hostfile support for POE and srun. MP_HOSTFILE env var to set
location of hostfile. Tasks will run from list order in the file.
-- Removes the slurmd's use of SysV shared memory. Instead the slurmd
communicates with the slurmstepd processes through the slurmstepd's
new named unix domain socket. The "stepd_api" is used to talk to the
slurmstepd (src/slurmd/common/stepd_api.[ch]).
-- Bluegene specific - bluegene block allocator will find most any
partition size now. Added support to start at any point in smap
to request a partition instead of always starting at 000.
-- Bluegene specific - Support to smap to down or bring up nodes in
configure mode. Added commands include allup, alldown,
up [range], down [range]
-- Time format in sinfo/squeue/smap/sacct changed from D:HH:MM:SS to
D-HH:MM:SS per POSIX standards document.
-- Treat scontrol update request without any requested changes as an
error condition.
-- Bluegene plugin renamed with BG instead of BGL. partition_allocator moved
into bluegene plugin and renamed block_allocator. Format for bluegene.conf
file changed also. Read bluegene html page. Code is backwards compatable
smap will generate in new form
-- Add srun option --nice to give user some control over job priority.
* Changes in SLURM 0.7.0-pre3
=============================
-- Restructure node states: DRAINING and DRAINED states are replaced
with a DRAIN flag. COMPLETING state is changed to a COMPLETING flag.
-- Test suite moved into testsuite/expect from separate repository.
-- Added new document describing slurm APIs (doc/html/api.html).
-- Permit nodes to be in multiple partitions simultaneously.
* Changes in SLURM 0.7.0-pre2
=============================
-- New stdio protocol. Now srun has just a single TCP stream to each node
of a job-step. srun and slurmd comminicate over the TCP stream using a
simple messaging protocol.
-- Added task plugin and use task prolog/epilog(s).
-- New slurmd_step functionality added. Fork exec instead of using shared
memory. Not completely tested.
-- BGL small partition logic in place in plugin and smap. Scheduler needs
to be rewritten to handle multiple partitions on a single node. No
documentation written on process yet.
-- If running select/bluegene plugin without access to BGL DB2, then
full-system bglblock is of system size defined in bluegene.conf.
* Changes in SLURM 0.7.0-pre1
=============================
-- Support defered initiation of job (e.g. srun --begin=11:30 ...).
-- Add support for srun --cpus-per-task through task allocation in
slurmctld.
-- fixed partition_allocator to work without curses
-- made change to srun to start message thread before other threads
to make sure localtime doesn't interfere.
-- Added new RPCs for slurmctld REQUEST_TERMINATE_JOB or TASKS,
REQUEST_KILL_JOB/TASKS changed to REQUEST_SIGNAL_JOB/TASKS.
-- Add support for e-mail notification on job state changes.
-- Some infrastructure added for task launch controls (slurm.conf:
TaskProlog, TaskEpilog, TaskPlugin; srun --task-prolog, --task-epilog).
* Changes in SLURM 0.6.11
=========================
-- Fix bug in sinfo partition sorting order.
-- Fix bugs in srun use of #SLURM options in batch script.
-- Use full Elan credential space rather than re-using credentials as soon
as job step completes (helps with fault-tolerance).
* Changes in SLURM 0.6.10
=========================
-- Fix for slurmd job termination logic (could hang in COMPLETING state).
-- Sacct bug fixes: Report correct user name for job step, show "uid.gid"
as fifth field of job step record (Andy Riebs, slurm.hp.sacct_uid.patch).
-- Add job_id to maui scheduler plugin start job status message.
-- Fix for srun's handling of null characters in stdout or stderr.
-- Update job accounting for larger systems (Andy Riebs, uptodate.patch).
Christopher J. Morrone
committed
-- Fixes for proctrack/linuxproc and mpich-gm support (Takao Hatazaki, HP).
-- Fix bug in switch/elan for large task count job having irregular task
distribution across nodes.
* Changes in SLURM 0.6.9
========================
-- Fix bug in mpi plugin to set the ID correctly
-- Accounting bug causing segv fixed (Andy Riebs, 14oct.jobacct.patch)
-- Fix for failed launch of a debugged job (e.g. bad executable name).
-- Wiki plugin fix for tracking allocated nodes (Ernest Artiaga, BSC).
-- Fix memory leaks in slurmctld and federation plugin.
-- Fix sefault in federation plugin function fed_libstate_clear().
-- Align job accounting data (Andy Riebs, slurm.hp.unal_jobacct.patch)
-- Restore switch state in backup controller restarts
* Changes in SLURM 0.6.8
========================
-- Invalid AllowGroup value in slurm.conf to not cause seg fault.
-- Fix bug that would cause slurmctld to seg-fault with select/cons_res
and batch job containing more than one step.
Moe Jette
committed
* Changes in SLURM 0.6.7
========================
-- Make proctrack/linuxproc thread safe, could cause slurmd seg fault.
-- Propagate umask from srun to spawned tasks.
-- Fix problem in switch/elan error handling that could hang a slurmd
step manager process.
-- Build on AIX with -bmaxdata:0x70000000 for memory limit more than 256MB.
-- Restore srun's return code support.
Moe Jette
committed
Moe Jette
committed
* Changes in SLURM 0.6.6
========================
-- Fix for bad socket close() in the spawn-io code.
Moe Jette
committed
* Changes in SLURM 0.6.5
========================
-- Sacct to report on job steps that never actually started.
-- Added proctrack/rms to elan rpm.
-- Restructure slurmctld/agent.c logic to insure timely reaping of
terminating pthreads.
-- Srun not to hang if job fails before task launches not all completed.
Moe Jette
committed
-- Fix for consumable resources properly scheduling nodes that have more
nodes than configured (Susanne Balle, HP, cons_res_patch.10.14.2005)
Moe Jette
committed
* Changes in SLURM 0.6.4
========================
-- Bluegene plugin drains an entire bglblock on repeated boot failures
only if it has not identified a specific node as being bad.
* Changes in SLURM 0.6.3
========================
-- Fix slurmctld mem leaks (step name and hostlist struct).
-- Bluegene plugin sets end time for job terminated due to removed
bglblock.
* Changes in SLURM 0.6.2
========================
-- Fix sinfo and squeue formatting to properly handle slurm nodes,
jobs, and other names containing "%".
* Changes in SLURM 0.6.1
========================
-- Fixed smap -Db to display slurm partitions correctly (take 2).
-- Add srun fork() retry logic for very heavily loaded system.
-- Fix possible srun hang on task launch failure.
-- Add support for mvapich v0.9.4, 0.9.5 and gen2.
* Changes in SLURM 0.6.0
========================
-- Add documentation for ProctrackType=proctrack/rms.
-- Make proctrack/rms be the default for switch/elan.
-- Do not preceed SIGKILL or SIGTERM to job step with (non-requested) SIGCONT.
-- Fixed smap -Db to display slurm partitions correctly.
-- Explicitly disallow ProctrackType=proctrack/linuxproc with
SwitchType=switch/elan. They will not work properly together.
* Changes in SLURM 0.6.0-pre8
=============================
-- Remove debugging xassert in switch/federation that were accidentally
committed
-- Make slurmd step manager retry slurm_container_destroy() indefinitely
instead of giving up after 30 seconds. If something prevents a job
step's processes from being killed, the job will be stuck in the
completing until the container destroy succeeds.
* Changes in SLURM 0.6.0-pre7
=============================
-- Disable localtime_r() calls from forked processes (semaphore set
in another pthread can deadlock calls to localtime_r made from
the forked process, this will be properly fixed in the next
major release of SLURM).
-- Added SLURM_LOCALID environment variable for spawned tasks
(Dan Palermo, HP).
-- Modify switch logic to restore state based exclusively upon
recovered job steps (not state save file).
-- Gracefully refuse job if there are too many job steps in slurmd.
-- Fix race condition in job completion that can leave nodes in
COMPLETING state after job is COMPLETED.
-- Added frees for BGL BrigeAPI strdups that were to this point unknown.
-- smap scrolls correctly for BGL systems.
-- slurm_pid2jobid() API call will now return the jobid for a step
manager slurmd process.
* Changes in SLURM 0.6.0-pre6
=============================
-- Added logic to return scheduled nodes to Maui scheduler (David
Jackson, Cluster Resources)
-- Fix bug in handling job request with maximum node count.
-- Fix node selection scheduling bug with heterogeneous nodes and
srun --cpus-per-task option
-- Generate error file to note prolog failures.
* Changes in SLURM 0.6.0-pre5
=============================
-- Modify sfree (BGL command) so that --all option no longer requires
an argument.
-- Modify smap so it shows all nodes and partitions by default (even
nodes that the user can't access, otherwise there are holes in
its maps).
-- Added module to parse time string (src/common/parse_time.c) for
future use.
-- Fix BlueGene hostlist processing for non-rectangular prisms and
add string length checking.
-- Modify orphan batch job time calculation for BGL to account for
slowness when booting many bglblocks at the same time.
* Changes in SLURM 0.6.0-pre4
=============================
-- Added etc/slurm.epilog.clean to kill processes initiated outside of
slurm when a user's last job on a node terminates.
-- Added config.xml and configurator.html files for use by OSCAR.
-- Increased maximum job step count from 64 to 130 for BGL systems only.
Christopher J. Morrone
committed
=============================
-- Add code so job request for shared nodes gets explicitly requested
nodes, but lightly loaded nodes otherwise.
-- Add job step name field.
-- Add job step network specification field.
-- Add proctrack/rms plugin
-- Change the proctrack API to send a slurmd_job_t pointer to both
slurm_container_create() and slurm_container_add(). One of those
functions MUST set job->cont_id.
-- Remove vestigial node_use (virtual or coprocessor) field from job
request RPC.
-- Fix mpich-gm bugs, thanks to Takao Hatazaki (HP).
-- Fix code for clean build with gcc 2.96, Takao Hatazaki (HP).
-- Add node update state of "RESUME" to return DRAINED, DRAINING, or
DOWN node to service (IDLE or ALLOCATED state).
-- smap keeps trying to connect to slurmctld in iterative mode rather
than just aborting on failure.
-- Add squeue option --node to filter by node name.
-- Modify squeue --user option to accept not only user names, but also
user IDs.
* Changes in SLURM 0.6.0-pre2
=============================
-- Removed "make rpm" target.
Christopher J. Morrone
committed
* Changes in SLURM 0.6.0-pre1
=============================
-- Added bgl/partition_allocator/smap changes from 0.5.7.
-- Added configurable resource limit propagation (Daniel Christians, HP).
-- Changed SlurmUser ID from 16-bit to 32-bit.
-- Added MpiDefault slurm.conf parameter.
-- Remove KillTree configuration parameter (replace with
"ProctrackType=proctrack/linuxproc")
-- Remove MpichGmDirectSupport configuration parameter (replace with
"MpiDefault=mpich-gm")
-- Make default plugin be "none" for mpi.
-- Added mpi/none plugin and made it the default.
-- Replace extern program_invocation_short_name with program_invocation_name
due to short name being truncated to 16 bytes on some systems.
-- Added support for Elan clusters with different CPU counts on nodes
(Chris Holmes, HP).
-- Added Consumable Resources web page (Susanne Balle, HP).
-- "Session manager" slurmd process has been eliminated.
-- switch/federation fixes migrated from 0.5.*
-- srun pthreads really set detached, fixes scaling problem
-- srun spawns message handler process so it can now be stopped (via
Ctrl-Z or TotalView) without inducing failures.
4417
4418
4419
4420
4421
4422
4423
4424
4425
4426
4427
4428
4429
4430
4431
4432
4433
4434
4435
4436
4437
4438
4439
4440
4441
4442
4443
4444
4445
4446
4447
4448
4449
4450
4451
4452
4453
4454
4455
4456
* Changes in SLURM 0.5.7
========================
-- added infrastructure for (eventual) support of AIX checkpointing
of slurm batch and interactive poe jobs
-- added wiring for BGL to do wiring for physical location first and then
logical.
-- only one thread used to query database before polling thread is there.
* Changes in SLURM 0.5.6
========================
-- fix for BGL hostnames and full system partition finding
* Changes in SLURM 0.5.5
========================
-- Increase SLURM_MESSAGE_TIMEOUT_MSEC_STATIC to 15000
-- Fix for premature timeout in _slurm_send_timeout
-- Fix for federation overlapping calls to non-thread-safe _get_adapters
* Changes in SLURM 0.5.4
========================
-- Added support for no reboot for VN to CO on BGL
-- Fix for if a job starts after it finishes on BGL
* Changes in SLURM 0.5.3
========================
-- federation patch so the slurm controller has sane window status at
start-up regardless of the window status reported in the slurmd
registration.
-- federation driver exits with fatal() if the federation driver can not
find all of the adapters listed in the federation.conf
* Changes in SLURM 0.5.2
========================
-- Extra federation driver sanity checks
* Changes in SLURM 0.5.1
========================
-- Fix federation driver bad free(), other minor fed fixes
-- Allow slurm to parse very long lines in the slurm.conf
* Changes in SLURM 0.5.0
========================
-- Fix race condition in job accouting plugin, could hang slurmd
-- Report SlurmUser id over 16 bits as an error (fix on v0.6)
* Changes in SLURM 0.5.0-pre19
==============================
-- Fix memory management bug in federation driver
* Changes in SLURM 0.5.0-pre18
==============================
-- elan switch plugin memory leak plugged
-- added g_slurmctld_jobacct_fini() to release all memory (useful
to confirm no memory leaks)
-- Fix slurmd bug introduced in pre17
* Changes in SLURM 0.5.0-pre17
==============================
-- slurmd calls the proctrack destroy function at job step completion
-- federation driver tries harder to clean up switch windows
Christopher J. Morrone
committed
* Changes in SLURM 0.5.0-pre16
==============================
-- Check slurm.conf values for under/overflows (some are 16 bit values).
Christopher J. Morrone
committed
-- Federation driver clears windows at job step completion
-- Modify code for clean build with gcc v4.0
Christopher J. Morrone
committed
-- New SLURM_NETWORK environmant variable used by slurm_ll_api
Christopher J. Morrone
committed
* Changes in SLURM 0.5.0-pre15
==============================
-- Added "network" field to "scontrol show job" output.
-- Federation fix for unfreed windows when multiple adapters on
one node use the same LID
* Changes in SLURM 0.5.0-pre14
==============================
-- RDMA works on fed plugin.
* Changes in SLURM 0.5.0-pre13
==============================
-- Major mods to support checkpoint on AIX.
-- Job accounting documenation expanded, added tuning options, minor bug fixes
-- BGL wiring will now work on <= 4 node X-dim partitions and also 8 node
X-dim partitions.
-- ENV variables set for spawning jobs.
-- jobacct patch from HP to not erroneously lock a mutex in the
jobacct_log plugin.
-- switch/federation supports multiple adapters per task. sn_all behaviour
is now correct, and it also supports sn_single.
* Changes in SLURM 0.5.0-pre12
==============================
-- Minor build changes to support RPM creation on AIX
* Changes in SLURM 0.5.0-pre11
==============================
-- Slurmd tests for initialized session manager (user's) slurmd pid before
killing it to avoid killing system daemon (race condition).
-- srun --output or --error file names of "none" mapped to /dev/null for
batch jobs rather than a file actually named "none".
-- BGL: don't try to read bglblock state until they are all created to
avoid having BGL Bridge API seg fault.
* Changes in SLURM 0.5.0-pre10
==============================
-- Fix bug that was resetting BGL job geometry on unrelated field update.
-- squeue and sinfo print timestamp in interate mode by default.
-- added scrolling windows in smap
-- introduced new variable to start polling thread in the bluegene plugin.
-- Latest accounting patches from Riebs/HP, retry communications.
-- Added srun option --kill-on-bad-exit from Holmes/HP.
-- Support large (64-bit address) log files where possible.
-- Fix problem of signals being delivered twice to tasks. Note that as
part of the fix the slurmd session manger no longer calls setsid to
create a new session.
* Changes in SLURM 0.5.0-pre9
=============================
-- If a job and node are in COMPLETING state and slurmd stops responding for
SlurmdTimeout, then set the node DOWN and the job COMPLETED.
-- Add logic to switch/elan to track contexts allocated to active job steps
rather than just using a cyclic counter and hoping to avoid collisions.
-- Plug memory leak in freeing job info retrieved using API.
-- Bluegene Plugin handles long deallocating states from driver 202.
-- Fix bug in bitfmt2int() which can go off allocated memory.
* Changes in SLURM 0.5.0-pre8
=============================
-- BlueGene srun --geometry was not getting propogated properly.
-- Fix race condition with multiple simultaneous epilogs.
-- Modify slurmd to resend job completion RPC to slurmctld in the
case where slurmctld is not responding.
-- Updated sacct: handle cancelled jobs correctly, add user/group
output, add ntasks ans synonym for nprocs, display error field
by default, display ncpus instead of nprocs
-- Parallelization of queing jobs up to 32 at once. Variable
MAX_AGENT_COUNT used in bgl_job_run.c to specify.
* Changes in SLURM 0.5.0-pre7
=============================
-- Preserve next_job_id across restarts.
-- Add support for really long job names (256 bytes).
-- Add configuration parameter SchedulerRootFilter to control what
entity manages prioritization of jobs in RootOnly partition
(internal scheduler plugin or external entity).
-- Added support for job accounting.
-- Added support for consumable resource based node scheduling.
-- Permit batch job to be launched to re-existing allocation.
* Changes in SLURM 0.5.0-pre6
=============================
-- Load bluegene.conf and federation.conf based upon SLURM_CONF env
var (if set).
-- Fix slurmd shutdown signal synchronization bug (not consistently
terminating).
-- Add doc/html/ibm.html document. Update bluegene.html.
-- Remove geometry[SYSTEM_DIMENSIONS] from opaque node_select data
type if SYSTEM_DIMENSIONS==0 (not ASCI-C compliant).
-- Modify smap to test for valid libdb2.so before issuing any BGL
Bridge API calls.
-- Modify spec file for optional inclusion of select_bluegene and
sched_wiki plugin libraries.
-- Initialize job->network in data structure, could cause job
submit/update to fail depending upon what is left on stack.
* Changes in SLURM 0.5.0-pre5
=============================
-- Expand buffer to hold node_select info in job termination log.
-- Modify slurmctld node hashing function to reduce collisions.
-- Treat bglblock vanishing as fatal error for job, prolog and epilog
exit immediately.
* Changes in SLURM 0.5.0-pre4
=============================
-- Fix bug in slurmd that could double KillWait time on job timeout.
-- Fix bug in srun's error code reporting to slurmctld, could DOWN
a node if job run as root has non-zero error code.
-- Remove a node's partition info when removed from existing partition.
-- Use proctrack plugin to call all processes in a job step before
calling interconnect_postfini() to insure no processes escape from
job and prevent switch windows from being released.
-- Added mail.html web page telling how to get on slurm mailing lists.
-- Added another directory to search for DB2 files on BGL system.
-- Added overview man page slurm.1.
-- Added new configure option "--with-db2-dir=PATH" for BGL.
* Changes in SLURM 0.5.0-pre3
=============================
-- Merge of SLURM v0.4-branch into v0.5/HEAD.
* Changes in SLURM 0.5.0-pre2
=============================
-- Fix bug in srun to clean-up upon failure of an allocated node
(srun -A would generate a segmentation fault, Chris Holmes, HP).
-- If slurmd's node name is mapped to NULL (due to bad configuration)
terminate slurmd with a fatal error and don't crash slurmctld.
-- Add SLURMD_DEBUG env var for use with AIX/POE in spawn_task RPC.
-- Always pack job's "features" for access by prolog/epilog
* Changes in SLURM 0.5.0-pre1
=============================
-- Add network option to srun and job creation API for specification
of communication protocol over IBM Federation switch.
-- Add new slurm.conf parameter ProctrackType (process tracking) and
associated plugin in the slurmd module.
-- Send node's switch state with job epilog completion RPC and
node registration (only when slurmd starts, not on periodic
registrtions).
-- Add federation switch plugin.
-- Add new configuration keyword, SchedulerRootFilter, to control
external scheduler control of RoolOnly partition (Chris Holmes, HP).
-- Modify logic to set process group ID for spawned processes (last
patch from slurm v0.3.11).
Moe Jette
committed
-- "srun -A" modified to return exit code of last command executed
(Chris Holmes, HP).
-- Add support for different slurm.conf files controlled via SLURM_CONF
env var (Brian O'Sullivan, pathscale)
-- Fix bug if srun given --uid without --gid option (Chris Holmes, HP).
4641
4642
4643
4644
4645
4646
4647
4648
4649
4650
4651
4652
4653
4654
4655
4656
4657
4658
4659
4660
4661
4662
4663
4664
4665
4666
4667
4668
4669
4670
4671
4672
4673
4674
4675
4676
4677
4678
4679
4680
4681
4682
4683
4684
4685
4686
4687
4688
4689
4690
4691
4692
4693
4694
4695
4696
4697
4698
4699
4700
4701
4702
4703
4704
4705
4706
4707
4708
4709
4710
4711
4712
4713
4714
4715
4716
4717
4718
4719
4720
4721
4722
4723
4724
4725
4726
4727
4728
4729
4730
4731
4732
4733
4734
4735
4736
4737
4738
4739
* Changes in SLURM 0.4.24
=========================
-- DRAIN nodes with switches on base partitions are in ERROR, MISSING,
or DOWN states.
* Changes in SLURM 0.4.23
=========================
-- Modified bluegene plugin to only sync bglblocks to jobs on initial
startup, not on reconfig. Fixes race condition.
-- Modified bluegene plugin to work with 141 driver. Enabling it to
only have to reboot when switching from coproc -> virtual and back.
-- added support for a full system partition to make sure every other
partition is free and vice-verse.
-- smap resizing issue fixed.
-- change prolog not to add time when a partition is in deallocating
state.
-- NOTE: This version of SLURM requires BGL driver 141/2005.
* Changes in SLURM 0.4.22
=========================
-- Modified bluegene plugin to not do anything if the bluegene.conf file
is altered.
-- added checking for lists before trying to create iterator on the list.
* Changes in SLURM 0.4.21
=========================
-- Fix in race condition with time in Status Thread of BGL
-- Fix no leading zeros in smap output.
* Changes in SLURM 0.4.20
=========================
-- Smap output is more user friendly with -c option
* Changes in SLURM 0.4.19
=========================
-- Added new RPCs for getting bglblock state info remotely and cache data
within the plugin (permits removal of DB2 access from BGL FEN and
dramatically increases smap responsivenss, also changed prolog/epilog
operation)
-- Move smap executable to main slurm RPM (from separate RPM).
-- smap uses RPC instead of DB2 to get info about bgl partitions.
-- Status function added to bluegene_agent thread. Keeps current state
of BGL partitions updating every second. will handle multiple attempts
at booting if booting a partition fails.
* Changes in SLURM 0.4.18
=========================
-- Added error checking of rm_remove_partition calls.
-- job_term() was terminating a job in real time rather than
queueing the request. This would result in slurmctld hanging
for many seconds when a job termination was required.
* Changes in SLURM 0.4.17
========================
-- Bug fixes from testing .16.
* Changes in SLURM 0.4.16
========================
-- Added error checking to a bunch of Bridge API calls and more
gracefully handle failure modes.
-- Made smap more robust for more jobs.
* Changes in SLURM 0.4.15
========================
-- Added error checking to a bunch of Bridge API calls and more
gracefully handle failure modes.
* Changes in SLURM 0.4.14
========================
-- job state is kept on warm start of slurm
* Changes in SLURM 0.4.13
========================
-- epilog fix for bgl plugin
* Changes in SLURM 0.4.12
========================
-- bug shot for new api calls.
-- added BridgeAPILogFile as an option for bluegene.conf file
* Changes in SLURM 0.4.11
========================
-- changed as many rm_get_partition() to rm_get_partitions_info as we could
for time saving.
* Changes in SLURM 0.4.10
========================
-- redesign for BGL external wiring.
-- smap display bug fix for smaller systems.
* Changes in SLURM 0.4.9
========================
-- setpnum works now, have to include this in bluegene.conf
* Changes in SLURM 0.4.8
========================
-- Changed the prolog and the epilog to use the env var MPIRUN_PARTITION
instead of BGL_PARTITION_ID
* Changes in SLURM 0.4.7
========================
-- Remove some BGL specific headers that IBM now distributes, NOTE
BGL driver 080 or greater required.
-- Change autogen.sh to deal with problems running autoconf on one
system and configure on another with different software versions.
-- took tv.h out of partition_allocator so it would work withn driver 080
from IBM.
-- updated slurmd signal handling to prevent possible user killing of daemon.
* Changes in SLURM 0.4.5
========================
-- Change sinfo default time limit field to have 10 bytes (up from 9).
-- Fix bug in bluegene partition selection (sorting bug).
-- Don't display any completed jobs in smap.
-- Add NodeCnt to filetxt job completion plugin.
-- Minor restructuring of how MMCS is polled for DOWN nodes and switches.
-- Fix squeue output format for "%s" (node select data).
-- Queue job requesting more resources than exist in a partition if
that partition's state is DOWN (rather than just abort it).
-- Add prolog/epilog for bluegene to code base (moved from mpirun in CVS)
-- Add prolog, epilog and bluegene.conf.example to bluegene RPM
-- In smap, Admin can get the Rack/midplane id from an XYZ input and vice versa.
-- Add smap line-oriented output capability.
* Changes in SLURM 0.4.4
========================
-- Fix race condition in slurmd seting pgid of spawned tasks for
process tracking.
-- Fix scontrol reconfig does nothing to running jobs nor crash the system
-- Fix sort of bgl_list only happens once in select_bluegene.c instead of every
time a new job is inserted.
* Changes in SLURM 0.4.3
========================
-- Turn off some RPM build checks (bug in RPM, see slurm.spec.in)
-- starting slurmctrld will destroy all RMP*** partitions everytime.
* Changes in SLURM 0.4.2
========================
-- Fix memory leak in BlueGene plugin.
-- Srun's --test-only option takes precedence over --batch option.
-- Add sleep(1) after setting bglblock owner due to apparent race condition
in the BGL API.
-- Slurm was timing out and killing batch jobs if the node registered when
a job prolog was still running.
-- BlueGene plugin kills jobs running in defunct bglblock on restart.
-- Smap displays pending jobs now, in addition to running and completing jobs.
-- Remove node "use=" from bluegene.conf file, create both coprocessor and
virtual bglblocks for now (later create just one and use API to change
it when such an API is available).
-- Add "ChangeNumpsets" parameter to bluegene.conf to use script to
update the numpsets parameter for newly created bglblocks (to be
removed once the API functions).
-- Add all patches from slurm v0.3.11 (through 2/7/2005)
- Added srun option --disable-status,-X to disable srun status feature
and instead forward SIGINT immediately to job upon receipt of Ctrl-C.
- Fix for bogus slurmd error message "Unable to put task N into pgrp..."
- Fix case where slurmd may erroneously detect shared memory entry
as "stale" and delete entry for unkillable or slow-to-exit job.
- (qsnet) Fix for running slurmd on node without and elan3 adapter.
- Fix for reported problem: slurm/538: user tasks block writing to stdio
-- Minor tweak to init.d/slurm for BlueGene systems.
-- Added smap RPM package (to install binary built on BlueGene
service node on front-end nodes).
-- Added wait between bglblock destroy and creation of new blocks
so that MMCS can complete the operation.
-- Fix bug in synchronizing bglblock owners on slurmctld restart.
* Changes in SLURM 0.4.0-pre11
==============================
-- Add new srun option "--test-only" for testing slurm_job_will_run API.
-- Fix bugs in slurm_job_will_run() processing.
-- Change slurm_job_will_run() to not return a message, just an error code.
-- Sync partition owners with running jobs on slurmctld restart.
* Changes in SLURM 0.4.0-pre10
==============================
-- Specify number of I/O nodes associated with BlueGene partition.
-- Do not launch a job's tasks if the job is cancelled while its
prolog is running (which can be slow on BlueGene).
-- Add new error code, ESLURM_BATCH_ONLY for attepts to launch
job steps on front-end system (e.g. Blue Gene).
-- Assorted fixes in smap, partition creation mode.
-- Add proper support for "srun -n" option on BGL recognizing
processor count in both virual and coprocessor modes.
-- Make default node_use on Blue Gene be coprocessor, as documented.
-- Add SIGKILL to BlueGene jobs as part of cleanup.
* Changes in SLURM 0.4.0-pre9
=============================
-- Change in /etc/init.d/slurm for RedHat and Suze compatability
* Changes in SLURM 0.4.0-pre8
=============================
-- Add logic to create and destroy Bluegene Blocks automatically as needed.
-- Update smap man page to include Bluegene configuration commands.
* Changes in SLURM 0.4.0-pre7
=============================
-- Port all patches from slurm v0.3 up through v0.3.10:
- Remove calls in auth/munge plugin deprecated by munge-0.4.
- Allow single task id to be selected with --input, --output, and --error.
- Create shared memory segment for Elan statistics when using the
switch/elan plugin.
- More fixes necessary for TotalView.
* Changes in SLURM 0.4.0-pre6
=============================
-- Add new job reason value "JobHeld" for jobs with priority==0
-- Move startup script from "/etc/rc.d/init.d/slurm" to "/etc/init.d/slurm"
-- Modify prolog/epilog logic in slurmd to accomodate very long run times,
on BGL these scripts wait for events that can take a very long time
(tens of seconds).
-- This code base was used for BGLb acceptance test with pre-defined
BGL blocks.
* Changes in SLURM 0.4.0-pre5
=============================
-- select/bluegene plugin confirms db.properties file in $sysconfdir
and copies it to StateSaveLocation (slurmctld's working directory)
-- select/bluegene plugin confirms environment variable required for
DB2 interaction are set (execute "db2profile" script before slurmctld)
-- slurmd to always give jobs KillWait time between SIGTERM and SIGKILL
at termination
-- set job's start_time and end_time = now rather than leaving zero if
they fail to execute
-- enable select/bluegene testing for DOWN nodes and switches
-- select/bluegene plugin to delete orphan jobs, free BGLblocks and
set owner as jobs terminate/start
* Changes in SLURM 0.4.0-pre4
=============================
-- Fixes for reported problems:
- slurm/512: Let job steps run on DRAINING nodes
- slurm/513: Gracefully deal with UIDs missing from passwd file
-- Add support for MPICH-GM (from takao.hatazaki@hp.com)
-- Add support for NodeHostname in node configuration
-- Make "scontrol show daemons" function properly on front-end system
(e.g. Blue Gene)
-- Fix srun bug when --input, --output and --error are all "none"
-- Don't schedule jobs for user root if partition is DOWN
-- Modify select/bluegene to honor job's required node list
-- Modify user name logic to explicitly set UID=0 to "root",
Suse Linux was not handling multiple users with UID=0 well.
* Changes in SLURM 0.4.0-pre3
=============================
-- Send SIGTERM to batch script before SIGKILL for mpirun cleanup on
Blue Gene/L
-- Create new allocation as needed for debugger in case old allocation
has been purged
-- Add Blue Gene User Guide to html documents
-- Fix srun bug that could cause seg fault with --no-shell option if not
running under a debugger
-- Propogate job's task count (if set) for batch job via SLURM_NPROCS.
-- Add new job parameters for Blue Gene: geometry, rotate, mode (virtual
or co-processor), communications type (mesh or torus), and partition ID.
-- Exercise a bunch of new switch plugin functions for Federation
switch support.
-- Fix bug in scheduling jobs when a processor count is specified
and FastSchedule=0 and the cluster is heterogeneous.
* Changes in SLURM 0.4.0-pre2
=============================
-- NOTE: "startclean" when transitioning from version 0.4.0-pre1, JOBS ARE LOST
-- Fixes for reported problems:
- slurm/477: Signal of batch job script (scancel -b) fixed
- slurm/481: Permit clearing of AllowGroups field for a partition
- slurm/482: Adjust Elan base context number to match RMS range
- slurm/489: Job completion logger was writing NULL to text file
-- Preserve job's requested processor count info after job is initiated
(for viewing by squeue and scontrol)
-- srun cancels created job if job step creation fails
-- Added a lots of Blue Gene/L support logic: slurmd executes on a single
node to front-end the 512-CPU base-partitions (Blue Gene/L's nodes)
-- Add node selection plugin infrastructure, relocate existing logic
to select/linear, add configuration parameter SelectType
-- Modify node hashing algorithm for better performance on Blue Gene/L
-- Add ability to specify node ranges for 3-D rectangular prism
* Changes in SLURM 0.4.0-pre1
=============================
-- NOTE: "startclean" when transitioning from version 0.3, JOBS ARE LOST
-- Added support for job account information (arbitrary string)
-- Added support for job dependencies (start job X after job Y completes)
-- Added support for configuration parameter CheckpointType
-- Added new job state "CANCELLED"
-- Don't strip binaries, breaks parallel debuggers
-- Fix bug in Munge authentication retry logic
-- Change srun handling of interupts to work properly with TotalView
-- Added "reason" field to job info showing why a job is waiting to run
* Changes in SLURM 0.3.7
========================
-- Fixes required for TotalView operability under RHEL3.0
(Reported by Dong Ahn <dahn@llnl.gov>)
- Do not create detached threads when running under parallel debugger.
- Handle EINTR from sigwait().
* Changes in SLURM 0.3.6
========================
-- Fixes for reported problems:
- slurm/459: Properly support partition's "Shared=force" configuration.
-- Resync node state to DRAINED or DRAINING on restart in case job
and node state recovered are out of sync.
-- Added jobcomp/script plugin (execute script on job completion,
from Nathan Huff, North Dakota State University).
-- Added new error code ESLURM_FRAGMENTED for immediate resource
allocation requests which are refused due to completing job (formerly
returned ESLURM_NOT_TOP_PRIORITY)
-- Modified job completion logging plugin calling sequence.
-- Added much of the infrastructure required for system checkpoint
(APIs, RPCs, and NULL plugin)
* Changes in SLURM 0.3.5
========================
-- Fix "SLURM_RLIMIT_* not found in environment" error message when
distributing large rlimit to jobs.
-- Add support for slurm_spawn() and associated APIs (needed for IBM
SP systems).
-- Fix bug in update of node state to DRAINING/DRAINED when update
request occurs prior to initial node registration.
-- Fix bug in purging of batch jobs (active batch jobs were being
improperly purged starting in version 0.3.0).
-- When updating a node state to DRAINING/DRAINED a Reason must be
provided. The user name and a timestamp will automatically be
appended to that Reason.
* Changes in SLURM 0.3.4
========================
-- Fixes for reported problems:
- slurm/404: Explicitly set pthread stack size to 1MB for srun
-- Allow srun to respond to ctrl-c and kill queued job while waiting
for allocation from controller.
* Changes in SLURM 0.3.3
========================
-- Fix slurmctld handling of heterogeneous processor count on elan
switch (was setting DRAINED nodes in state DRAINING).
-- Fix sinfo -R, --list-reasons to list all relevant node states.
-- Fix slurmctld to honor srun's node configuration specifications
with FastSchedule==0 configuration.
-- Added srun option --debugger-test to confirm that slurm's debugger
infrastructure is operational.
-- Removed debugging hacks for srun.wrapper.c. Temporarily use
RPM's debugedit utility if available for similar effect.
* Changes in SLURM 0.3.2