Monitoring & Reliability October 2026

How to Get Email Alerts When Your Python Cron Job Fails

Silent cron job failures are the most expensive bugs in production. Here's how to set up email alerts so you know immediately when a Python scheduled job fails, crashes, or stops running.

L
LiteLambda Team
7 min read

If your Python cron job fails at 2am, do you find out at 2am — or when a customer complains at 9am?

Most developers don't find out until something downstream breaks. A scheduled report doesn't arrive. A payment retry queue stops processing. A database sync falls behind. By the time the symptom surfaces, the job has been silently failing for hours.

This guide shows you exactly how to set up email alerts for Python cron job failures — so you know the moment something stops working, not hours later.


Why Python Cron Jobs Fail Silently

The classic crontab has no built-in failure notification. When a job exits with a non-zero status code, cron does nothing. When a job runs fine but produces wrong output, cron does nothing. When a job silently hangs past its scheduled time, cron does nothing.

There are four ways a cron job can fail without you knowing:

  1. Exception during execution — requests.get() raises a ConnectionError, the script exits with code 1. Cron moves on.
  2. Timeout / hang — Your job is waiting on a database lock or slow API. It runs forever. The next invocation starts alongside it.
  3. Silent logic failure — The job runs and exits 0, but processed 0 records instead of 500. No error. No alert.
  4. Job never starts — Server reboot wiped the crontab. Or the system clock drifted. Or the venv broke. The job simply doesn't run.

Cases 1–3 are caught by exception monitoring. Case 4 — the job never starting — requires a "dead man's switch" pattern (more on that below).


Option 1: Email Alerts via Try/Except + SMTP (Manual Approach)

The most portable solution wraps your job in a try/except block and sends an email on failure:

import smtplib
import traceback
from email.mime.text import MIMEText
from email.mime.multipart import MIMEMultipart

ALERT_EMAIL = "[email protected]"
SMTP_HOST = "smtp.gmail.com"
SMTP_PORT = 587
SMTP_USER = "[email protected]"
SMTP_PASS = "your-app-password"

def send_failure_alert(job_name: str, error: Exception):
    msg = MIMEMultipart()
    msg["Subject"] = f"🚨 Cron Job Failed: {job_name}"
    msg["From"] = SMTP_USER
    msg["To"] = ALERT_EMAIL

    body = f"""
Your scheduled Python job <strong>{job_name}</strong> has failed.

<b>Error:</b> {type(error).__name__}: {str(error)}

<b>Traceback:</b>
<pre>{traceback.format_exc()}</pre>

Time: {__import__('datetime').datetime.utcnow().isoformat()}
    """
    msg.attach(MIMEText(body, "html"))

    with smtplib.SMTP(SMTP_HOST, SMTP_PORT) as server:
        server.starttls()
        server.login(SMTP_USER, SMTP_PASS)
        server.send_message(msg)


def run_my_job():
    # Your actual job logic here
    import requests
    response = requests.get("https://api.example.com/data", timeout=30)
    response.raise_for_status()
    data = response.json()
    print(f"Processed {len(data)} records")


if __name__ == "__main__":
    try:
        run_my_job()
    except Exception as e:
        send_failure_alert("my-daily-sync", e)
        raise  # Re-raise so cron gets a non-zero exit code

Pros: Works anywhere with an SMTP server.
Cons: You're maintaining SMTP config in every script, and this doesn't catch Case 4 (job never runs).


Option 2: Failure Alerts via a Monitoring Platform

Dedicated monitoring handles all four failure cases automatically. The pattern:

  1. Your job runs and sends a "heartbeat" ping to the monitoring service when it completes successfully
  2. If the monitoring service doesn't receive a heartbeat within the expected window, it emails you

This is called a dead man's switch or heartbeat monitoring.

import requests
import os

HEARTBEAT_URL = os.environ.get("HEARTBEAT_URL")  # From your monitoring service

def run_my_job():
    # Your job logic
    process_payments()
    sync_inventory()

    # Signal success
    if HEARTBEAT_URL:
        requests.get(HEARTBEAT_URL, timeout=5)

if __name__ == "__main__":
    run_my_job()

If run_my_job() raises an exception, the heartbeat ping never happens → monitoring service alerts you.

If the job never starts at all (Case 4), the heartbeat never pings → monitoring service alerts you.


If you're running Python cron jobs on LiteLambda, failure alerts are built in — no SMTP setup, no heartbeat URLs, no extra monitoring service to configure.

How it works:

  1. Write your Python handler function:
import requests

def handler(event, context):
    # Your job runs here
    response = requests.get("https://api.example.com/data", timeout=30)
    response.raise_for_status()

    records = response.json()
    processed = 0

    for record in records:
        process_record(record)
        processed += 1

    print(f"✓ Processed {processed} records")
    return {"processed": processed}
  1. In the LiteLambda dashboard, enable Failure Notifications for the job:
  2. Set your alert email address
  3. Choose: alert on exception, alert on non-zero exit, alert on timeout, or alert if job doesn't run

  4. Done. If your job crashes, times out, or misses its scheduled window, you get an email with:

  5. The job name and cron expression
  6. The full Python traceback (if an exception occurred)
  7. The exact timestamp of failure
  8. A direct link to the execution log

Deploy with the CLI:

pip install litelambda-cli
litelambda login
litelambda deploy my-daily-sync handler.py --cron "0 2 * * *" --alert-email [email protected]

This handles all four failure cases automatically, including jobs that never run (Case 4) — which SMTP-based solutions can't catch.


Comparison: Three Approaches

Approach Setup Time Catches All Failure Types? Cost
Try/except + SMTP ~30 min ❌ Won't catch "job never ran" SMTP costs
Healthchecks.io + heartbeat ~20 min ✅ Yes Free tier available
LiteLambda built-in alerts ~5 min ✅ Yes (including "never ran") Included in plan

Setting Up Alerts for Existing VPS/Server Cron Jobs

If you're on a VPS and can't migrate the job yet, use the heartbeat pattern with a ping URL:

# In your crontab (replace with your actual command and ping URL)
0 2 * * * /usr/bin/python3 /home/user/my_script.py && curl -fsS --retry 3 https://hc-ping.example.com/your-uuid > /dev/null

The && means the curl only runs if the Python script exits 0. If the script fails, no ping → alert fires.

For more complex failure modes (the script exits 0 but processed 0 records), add explicit validation:

def handler(event, context):
    records = fetch_records()

    if len(records) == 0:
        raise ValueError("Expected records to process, got 0. Possible upstream issue.")

    for record in records:
        process_record(record)

    return {"processed": len(records)}

Never silently swallow empty results — make your job fail loudly if the data doesn't look right.


What Your Alert Email Should Include

A useful failure alert contains:

  • Job name and schedule — Which job, and when was it supposed to run?
  • Full traceback — The actual Python exception, not just "job failed"
  • Execution duration — Did it fail immediately (auth issue) or after 4 minutes (timeout)?
  • Previous successful run — How long has it been working normally?
  • Link to full logs — For digging into output before the crash

LiteLambda's built-in alerts include all of these. SMTP-based solutions require you to build this manually.


Common Mistakes with Cron Job Alerting

Mistake 1: Only alerting on exceptions, not on missed runs.
A job that never starts is worse than one that fails and alerts. Always use heartbeat monitoring alongside exception alerts.

Mistake 2: Using except Exception: pass.
This silently swallows all errors. Always re-raise or log explicitly:

try:
    run_job()
except Exception as e:
    logger.error(f"Job failed: {e}", exc_info=True)
    raise  # Don't swallow it

Mistake 3: Alerting on every retry.
If your job retries 3 times before giving up, only alert after all retries are exhausted:

MAX_RETRIES = 3

def handler(event, context):
    for attempt in range(MAX_RETRIES):
        try:
            run_job()
            return {"status": "success"}
        except Exception as e:
            if attempt == MAX_RETRIES - 1:
                raise  # Final attempt — let the alert fire
            print(f"Attempt {attempt + 1} failed, retrying...")

Mistake 4: No alert on slow jobs.
A job that runs for 4 hours instead of 5 minutes is a problem even if it eventually succeeds. Set a timeout and alert on it.


Summary

Silent cron job failures are solved with one of three approaches:

  • SMTP try/except — Works, but doesn't catch missed runs and requires per-script setup
  • Heartbeat monitoring — Catches all failure types, requires instrumenting every job
  • LiteLambda built-in alerts — Catches all failure types with zero instrumentation, ~5 min setup

If you're already running Python scheduled jobs and want to add failure alerts in the next 10 minutes without changing your code: deploy your script to LiteLambda and enable notifications from the dashboard. Your existing Python code works as-is — wrap it in a handler(event, context) function and you're done.


Related: How to Monitor Python Cron Jobs in 2026 · Python Dead Man's Switch: Alert If Your Cron Doesn't Run · Heroku Scheduler Alternative for Python Developers

Skip the infrastructure setup.

Run this exact code in our secure, isolated Docker sandbox. It takes 10 seconds to deploy.

Deploy this script in 60s →

No DevOps required.