Lease renewal batch job stuck in processing state, no error logs generated

Our scheduled lease renewal batch job has been stuck in ‘PROCESSING’ state for over 18 hours with no progress. The job normally completes in 2-3 hours for our monthly renewal cycle. The scheduler shows it as running, but CPU and memory usage for the job process are near zero, suggesting it’s not actually doing anything.

What’s concerning is that there are no error messages in any of the standard log files - application log, scheduler log, or database log all show the job started successfully but nothing after that. The job status in the batch job monitor just shows ‘PROCESSING - 0% complete’ with the start timestamp from yesterday morning.

I’m hesitant to kill the job forcefully because we have about 340 leases queued for renewal processing and I’m worried about data corruption if records are mid-update. Has anyone experienced batch jobs getting stuck like this without generating any error output? What’s the safest way to diagnose what’s blocking it?

Your batch job is stuck due to a combination of factors involving the scheduler, locked records, and insufficient debug logging. Let me address each focus area:

Batch Job Scheduler: The ICS 2023-1 batch scheduler has a known issue where jobs can enter a zombie state if the scheduler service experiences a brief network interruption or memory pressure event. The job process continues running, but it loses its connection to the scheduler coordination service. This explains why you see ‘PROCESSING’ status with 0% progress and near-zero resource usage.

To diagnose scheduler connectivity:

  1. Check scheduler service health: Navigate to Admin → System Monitor → Scheduler Service Status
  2. Look for ‘Orphaned Jobs’ count - if non-zero, you have disconnected job processes
  3. Review scheduler service logs at /logs/scheduler-service.log for connection reset events around your job start time

The safe recovery procedure:

  1. Mark the job as failed in the scheduler: `scheduler-admin --job-id --force-fail
  2. This releases the job from the scheduler’s active queue without killing the process
  3. The job process will detect the status change and terminate gracefully within 5 minutes
  4. Any acquired locks will be released through normal transaction rollback

Locked Records: Your observation about 12 database sessions with locks belonging to the job itself indicates the job successfully acquired locks on the first batch of lease records but then stalled before processing them. This is characteristic of a resource deadlock or external dependency failure.

To identify the blocking resource:

SELECT s.sid, s.serial#, s.wait_class, s.event, s.seconds_in_wait,
       l.object_id, o.object_name, l.locked_mode
FROM v$session s
JOIN v$lock l ON s.sid = l.sid
JOIN dba_objects o ON l.object_id = o.object_id
WHERE s.program LIKE '%LeaseRenewal%'
ORDER BY s.seconds_in_wait DESC;

If the wait_event shows ‘SQLNet message from client’ or 'SQLNet more data from client’, the job is waiting for application-layer processing, not database resources. This suggests the batch job code has a deadlock or infinite loop.

Safe lock release procedure:

  1. Identify the session IDs from the query above
  2. For each session, check if it has uncommitted work: `SELECT USED_UREC FROM V$TRANSACTION WHERE ADDR IN (SELECT TADDR FROM V$SESSION WHERE SID = )
  3. If USED_UREC > 0, the session has uncommitted changes that will be rolled back
  4. Kill the sessions: `ALTER SYSTEM KILL SESSION ‘,<serial#>’ IMMEDIATE;
  5. Monitor rollback progress: Check V$FAST_START_TRANSACTIONS for recovery status

Data corruption risk is minimal because ICS batch jobs use ACID-compliant transactions. When you kill the sessions, Oracle will automatically roll back any uncommitted work, leaving your lease data in a consistent state.

Debug Logging: The absence of error logs indicates your batch job logging configuration is insufficient for troubleshooting. ICS 2023-1 batch jobs have multiple logging levels that must be enabled separately.

Enable comprehensive batch job logging:

  1. Edit batch-job-config.properties:

batch.job.logging.level=DEBUG
batch.job.progress.reporting=true
batch.job.checkpoint.logging=true
batch.job.transaction.trace=true
batch.job.exception.stacktrace=true
  1. Enable lease-specific debug logging:

lease.renewal.debug=true
lease.renewal.record.level.trace=true
  1. Configure separate log file for batch jobs:

batch.job.log.file=/logs/batch-jobs/lease-renewal.log
batch.job.log.rotation=daily
batch.job.log.retention=30
  1. Restart the batch job service to apply changes

With debug logging enabled, you’ll see detailed output including:

  • Each lease record being processed
  • Transaction commit/rollback points
  • External API calls (if any)
  • Lock acquisition and release events
  • Progress percentage calculations
  • Any exceptions or warnings, even if they’re caught and handled

For immediate diagnosis of your current stuck job, you can enable runtime logging without restart:


scheduler-admin --job-id <ID> --set-log-level DEBUG

This will start generating logs for the running job (if it’s actually processing), or confirm it’s completely stalled (if no new log entries appear).

Recommended Resolution:

  1. Enable debug logging configuration first (for future jobs)
  2. Use the scheduler force-fail command to safely terminate the stuck job
  3. Wait for database rollback to complete (monitor V$FAST_START_TRANSACTIONS)
  4. Verify all locks are released: Check v$lock for your lease tables
  5. Investigate scheduler service logs for the root cause of the disconnect
  6. Restart the lease renewal job with debug logging enabled
  7. Monitor the new job execution with detailed logs to ensure it progresses normally

The underlying issue is likely scheduler service instability. Check for memory pressure, network issues, or service restarts around the time your job started. Consider implementing job timeouts (recommended: 6 hours for your 340-lease workload) to prevent indefinite stuck states in the future.


This draft is based on general Infor CloudSuite knowledge. It has not been verified against your specific version and environment. Practitioners: verify the steps and share your experience below.

This sounds like a database lock issue. The job is probably waiting for a lock on the lease records that’s being held by another process or transaction. Check for long-running database sessions or uncommitted transactions. Query your database for locks on the Lease table and see if there’s a blocking session. The lack of error logs makes sense because the job hasn’t actually failed - it’s just waiting indefinitely for a resource.

Good call on the database locks. I found 12 active sessions with locks on lease records, but they all belong to the batch job itself. It looks like the job acquired locks on the first batch of leases and then got stuck before processing them. The locks are now 19+ hours old. Should I kill the database sessions to release the locks, or will that corrupt the lease data?

Before killing anything, check if the batch job is actually deadlocked or just waiting on a specific resource. Look at the wait events for those database sessions - if they’re waiting on ‘enq: TX - row lock contention’, you have a deadlock situation. If they’re waiting on something else like network I/O or file system access, the problem might be external to the database. Also check if your batch job logging is configured correctly - sometimes jobs appear to run silently because the log level is set too high or the log file path is inaccessible.

I’ve seen this exact behavior when the batch job framework loses connection to the scheduler service but doesn’t detect the disconnect. The job thinks it’s running, the scheduler thinks it’s running, but they’re not actually communicating. Check the scheduler service health and restart it if necessary. The job should detect the scheduler disconnect and either fail gracefully or resume. Also verify that your batch job has proper timeout configuration - a job with no timeout can run indefinitely in a stuck state.

Another possibility: check if your lease renewal job is configured with proper transaction boundaries. If it’s processing all 340 leases in a single transaction without intermediate commits, it could appear stuck while actually processing very slowly due to transaction log growth and lock escalation. The 0% complete status suggests the job might not be reporting progress correctly even if it’s working.

Check the job’s checkpoint configuration. If checkpointing is disabled, the job won’t report progress and you’ll see 0% throughout execution.