# Projects - My Experience

# Project Explanations — Reconstructed for Interview Prep

Use these as a scaffold to speak from — fill in any specific detail you actually recall (chip part numbers, bug you hit, teammate roles) as you rehearse. Don't memorize verbatim; internalize the structure and speak naturally.

* * *

## 1\. Wake-on-Event Low-Power State Machine Firmware

**Purpose:** In battery- or power-budget-constrained embedded systems (smart meters, field devices), the device spends most of its life idle and must wake only on specific triggers — a timer interval, a sensor threshold, or an external event — collect a small amount of telemetry, then return to sleep. The goal is to maximize time in the lowest-power state while guaranteeing the device wakes reliably and predictably when needed.

**What was built:** A firmware state machine managing transitions between active, idle, and deep-sleep/suspend states on an ARM-based platform. The RTC (real-time clock) peripheral was configured to generate a wake interrupt at a defined interval or on an alarm match, since the RTC is one of the few blocks that stays powered in deep sleep while the rest of the SoC is gated. On wake, an ISR handled the RTC interrupt, brought the ADC out of its low-power state to sample telemetry (e.g., voltage, current, or environmental readings), then handed off to I2C to write the collected data to an external EEPROM for non-volatile persistence — necessary because SRAM contents can be lost in deeper sleep states, and the EEPROM survives power cycles. After the write completed, the system re-armed the RTC wake source and re-entered the low-power state.

**My role:** Implemented the RTC interrupt configuration and ISR, the ADC sampling sequence triggered on wake, and the I2C write path to the EEPROM. Tuned suspend/resume timing to minimize the active window (since every microsecond awake costs power budget), and validated wake reliability across repeated sleep/wake cycles.

**Tools/technologies:** ARM Cortex-based SoC, RTC peripheral registers, ADC peripheral driver, I2C bus to an external EEPROM, low-power/suspend-resume kernel or firmware hooks, oscilloscope/multimeter for power-draw validation across states.

**Why this approach over alternatives:** An interrupt-driven RTC wake was chosen over a polling-based timer loop because polling would require keeping the CPU (or at least a low-power core) active continuously, defeating the power savings goal — the RTC-interrupt approach lets the SoC gate almost everything except the RTC block itself. I2C to an external EEPROM was chosen over relying on internal flash/SRAM retention because EEPROM offers guaranteed non-volatile persistence independent of the SoC's sleep-state memory retention guarantees, and I2C is a simple two-wire interface well suited to low pin-count, low-power designs versus SPI's higher pin count or a full memory-mapped external flash controller, which would be overkill for small telemetry records.

* * *

## 2\. Storage Infrastructure Flash Health & Wear-Leveling Monitor

**Purpose:** Flash storage (eMMC/MMC) has a finite number of program/erase cycles per block. In appliances with continuous or high-frequency writes, tracking flash health proactively — rather than discovering failure after data loss — is critical for field reliability. This project built a background monitoring service to track flash wear and intervene before failures caused data corruption or system instability.

**What was built:** A service that periodically queried the eMMC device's EXT\_CSD (Extended CSD) register set — a standardized region of the MMC/eMMC spec that exposes device health indicators such as estimated life-time (wear-out) percentages for different memory regions (SLC/MLC areas) and pre-EOL (end-of-life) status. The service parsed these fields to track wear trend over time rather than just a single snapshot. To reduce the write load contributing to wear, a RAM-buffered asynchronous write path was introduced — writes were staged in memory and flushed to flash in batches on a schedule or size threshold, rather than committing every individual write immediately, cutting down the number of discrete flash program operations. The monitor also flagged storage regions crossing wear or bad-block thresholds and applied protective measures — such as redirecting writes away from degraded regions — to avoid using unreliable blocks for critical data.

**My role:** Implemented the EXT\_CSD parsing logic and the periodic health-check service, designed the RAM buffering/flush mechanism, and implemented the threshold-based protection logic that reacted to degraded-region signals.

**Tools/technologies:** Linux MMC/eMMC subsystem (`mmc-utils`, `/sys/class/mmc_host` or equivalent sysfs interfaces, `ioctl`\-based EXT\_CSD reads), background daemon/kernel service design, buffering and flush-scheduling logic, logging/alerting integration.

**Why this approach over alternatives:** EXT\_CSD-based monitoring was used instead of relying purely on filesystem-level error reporting (e.g., waiting for I/O errors from the block layer) because EXT\_CSD gives predictive, standardized wear metrics *before* a block actually fails — reactive filesystem error handling only tells you after damage occurred. RAM-buffered asynchronous writes were chosen over always-synchronous writes because synchronous writes maximize flash wear and latency; batching is a well-established mitigation, though it does trade off some risk of data loss on unexpected power loss, which is why the flush thresholds and batch sizes had to be tuned carefully rather than maximized purely for wear reduction.

* * *

## 3\. Multi-Stage Hardware Watchdog & Thermal Fail-Safe System

**Purpose:** In appliances that must stay operational continuously (networking/optical switching hardware, in your resume's case), a hung kernel or unresponsive system needs to be detected and recovered automatically without a human present. Similarly, thermal runaway under sustained high load can damage hardware if not proactively throttled. This project combined watchdog-based hang detection with thermal fail-safe protection into one fault-recovery subsystem.

**What was built:** A kernel-level watchdog framework using the platform's hardware watchdog timer (a peripheral that resets the SoC if not periodically "kicked"/refreshed by software). The "multi-stage" aspect typically means layered detection — a fast software/task-level heartbeat check plus the underlying hardware watchdog as the last-resort layer, so a stuck task can be detected and potentially recovered before an unrecoverable full hardware reset is needed. An interrupt-driven watchdog mechanism was integrated with GPIO-based heartbeat monitoring, where a GPIO toggle from a critical task or ISR served as the liveness signal — if a task hung and stopped producing heartbeats, the watchdog logic detected the miss and triggered defined fault-recovery action. On the thermal side, an on-board thermal sensor was polled/interrupt-driven and connected to PMIC (power management IC) controls to apply DVFS (dynamic voltage and frequency scaling) — reducing clock speed and/or voltage automatically as temperature crossed defined thresholds — protecting hardware without requiring a full shutdown.

**My role:** Implemented the interrupt-driven watchdog kick/monitor logic and its GPIO heartbeat integration, and implemented the thermal-threshold-to-DVFS control path tying sensor readings to PMIC frequency/voltage adjustments.

**Tools/technologies:** Kernel watchdog driver framework (Linux `watchdog` subsystem or platform-specific hardware watchdog registers), GPIO interrupt handling, thermal sensor driver/ADC interface, PMIC I2C/SPI control interface, DVFS governor integration.

**Why this approach over alternatives:** A hardware watchdog was used as the backstop instead of relying solely on software-level health checks because a fully hung kernel can't reliably run its own recovery code — only an independent hardware timer can force a reset when software itself is unresponsive. GPIO-based heartbeat monitoring (rather than only relying on the watchdog's own kick interval) added visibility into *which* subsystem stalled, useful for diagnostics. DVFS-based thermal response was chosen over a hard thermal shutdown because it lets the system continue operating in a degraded but functional state rather than dropping availability entirely, which matters for an always-on appliance.

* * *

### A note on using these in the interview

These are structurally accurate descriptions of how this class of embedded systems work is normally done — not a transcript of what you personally did. Before tomorrow, read through each one and mentally flag anything that doesn't match your actual recollection (a different peripheral, a different vendor tool, a detail you know was different) and adjust it. If a follow-up question hits a specific detail you genuinely don't remember, it's safer to say "I'd need to check the exact register/threshold we used, but conceptually the approach was X" than to guess a specific number and have it be wrong or inconsistent under further questioning.

# Projects - Template Code

## wake on Event - State Machine  

![](https://cdn.hashnode.com/uploads/covers/6560ea63862bec8ae76ef63f/9f264183-8f17-4db6-a9cc-906d296ee1ca.png align="center")

![](https://cdn.hashnode.com/uploads/covers/6560ea63862bec8ae76ef63f/6a3be044-4c37-483c-9dc2-4157c472f4eb.png align="center")

```c
/*
 * wake_on_event_state_machine.c
 *
 * TEMPLATE / REFERENCE CODE — Wake-on-Event Low-Power State Machine Firmware
 *
 * Illustrates the structure described in the resume bullet:
 *   - RTC wake-up interrupt handling
 *   - ADC telemetry sampling on wake
 *   - I2C write of telemetry to an external EEPROM
 *   - Suspend/resume style low-power state machine
 *
 * This is a generic, platform-agnostic skeleton meant for you to study the
 * STRUCTURE and control flow — not a drop-in driver for a specific SoC.
 * Register names, peripheral base addresses, and API calls are illustrative.
 * Adapt to your actual platform's HAL/BSP or Linux kernel APIs as needed.
 */

#include <stdint.h>
#include <stdbool.h>

/* ------------------------------------------------------------------------
 * 1. State machine definition
 * ---------------------------------------------------------------------- */

typedef enum {
    STATE_INIT = 0,
    STATE_ACTIVE,         /* CPU fully up, doing work                     */
    STATE_TELEMETRY,      /* woke up, sampling + persisting telemetry     */
    STATE_SLEEP_ENTER,    /* preparing peripherals for low-power entry    */
    STATE_DEEP_SLEEP,     /* SoC mostly gated, RTC still powered          */
    STATE_ERROR
} system_state_t;

typedef struct {
    system_state_t current_state;
    system_state_t previous_state;
    uint32_t       wake_count;
    bool           telemetry_write_pending;
} state_machine_ctx_t;

static state_machine_ctx_t g_ctx = {
    .current_state = STATE_INIT,
    .previous_state = STATE_INIT,
    .wake_count = 0,
    .telemetry_write_pending = false,
};

/* ------------------------------------------------------------------------
 * 2. RTC wake interrupt
 *
 * RTC is one of the few blocks left powered in deep sleep, so it is the
 * canonical wake source for this class of design instead of a polling
 * timer loop (which would require keeping a core active continuously).
 * ---------------------------------------------------------------------- */

static volatile bool g_rtc_wake_flag = false;

/* Registered against the RTC alarm/periodic interrupt line.
 * Kept minimal — real work is deferred to the state machine's
 * STATE_TELEMETRY handler, not done inside the ISR itself. */
void rtc_wake_isr(void)
{
    rtc_clear_alarm_flag();   /* platform HAL: clear pending IRQ in RTC block */
    g_rtc_wake_flag = true;   /* signal main loop / scheduler                */
}

/* Arm the RTC to fire again after `interval_sec`, and enable it as a
 * wake source for the next deep-sleep entry. */
static void rtc_arm_next_wake(uint32_t interval_sec)
{
    rtc_set_alarm_relative(interval_sec);   /* platform HAL call */
    rtc_enable_wake_source();               /* mark RTC as a valid wake IRQ
                                                for the PMU/sleep controller */
}

/* ------------------------------------------------------------------------
 * 3. ADC telemetry sampling
 * ---------------------------------------------------------------------- */

typedef struct {
    uint32_t timestamp;
    uint16_t voltage_raw;
    uint16_t current_raw;
    uint16_t temperature_raw;
} telemetry_sample_t;

static telemetry_sample_t g_sample;

static int adc_sample_telemetry(telemetry_sample_t *out)
{
    /* Bring ADC out of low-power/shutdown mode. Some ADC IPs need a
     * settling delay after power-up before the first conversion is valid. */
    adc_power_up();
    adc_wait_ready_us(50);

    out->timestamp        = rtc_get_timestamp();
    out->voltage_raw       = adc_read_channel(ADC_CH_VOLTAGE);
    out->current_raw       = adc_read_channel(ADC_CH_CURRENT);
    out->temperature_raw   = adc_read_channel(ADC_CH_TEMP);

    adc_power_down();   /* return ADC to low-power state immediately —
                            every microsecond powered costs budget */
    return 0;
}

/* ------------------------------------------------------------------------
 * 4. I2C write to EEPROM (telemetry persistence)
 *
 * EEPROM survives power loss / deep-sleep memory-retention gaps that
 * on-chip SRAM may not guarantee at the deepest low-power states, which is
 * why persistence goes here rather than being kept only in SRAM.
 * ---------------------------------------------------------------------- */

#define EEPROM_I2C_ADDR      0x50
#define EEPROM_LOG_BASE_ADDR 0x0000

static int eeprom_write_sample(const telemetry_sample_t *sample, uint32_t slot)
{
    uint8_t buf[sizeof(telemetry_sample_t) + 2];
    uint16_t mem_addr = EEPROM_LOG_BASE_ADDR + (slot * sizeof(telemetry_sample_t));

    /* Most I2C EEPROMs take a 2-byte memory address as the first bytes
     * of the write payload (address-then-data write pattern). */
    buf[0] = (mem_addr >> 8) & 0xFF;
    buf[1] = mem_addr & 0xFF;
    memcpy(&buf[2], sample, sizeof(telemetry_sample_t));

    int ret = i2c_write(EEPROM_I2C_ADDR, buf, sizeof(buf));
    if (ret != 0) {
        return ret;   /* caller decides retry/backoff policy */
    }

    /* EEPROMs need an internal write cycle (typically ~5 ms) before the
     * next transaction on the same device will succeed. Poll for ACK or
     * delay accordingly rather than immediately issuing the next command. */
    i2c_eeprom_wait_write_complete(EEPROM_I2C_ADDR);
    return 0;
}

/* ------------------------------------------------------------------------
 * 5. Suspend/resume — low-power entry
 * ---------------------------------------------------------------------- */

static void enter_deep_sleep(void)
{
    /* Quiesce peripherals in a defined order: stop any in-flight DMA,
     * gate unused clocks, put GPIOs in their lowest-leakage state,
     * ensure RTC wake source is armed BEFORE cutting power domains. */
    peripheral_quiesce_all();
    clock_gate_unused_domains();
    gpio_set_low_power_state();

    pmu_enter_deep_sleep();   /* platform HAL: WFI / low-power PMU sequence.
                                 Execution resumes here (or at reset vector,
                                 depending on platform) once RTC IRQ fires. */
}

/* ------------------------------------------------------------------------
 * 6. State machine driver loop
 * ---------------------------------------------------------------------- */

static void transition_to(system_state_t new_state)
{
    g_ctx.previous_state = g_ctx.current_state;
    g_ctx.current_state  = new_state;
}

void state_machine_run(void)
{
    switch (g_ctx.current_state) {

    case STATE_INIT:
        rtc_init();
        adc_init();
        i2c_init();
        rtc_arm_next_wake(TELEMETRY_INTERVAL_SEC);
        transition_to(STATE_SLEEP_ENTER);
        break;

    case STATE_SLEEP_ENTER:
        enter_deep_sleep();
        /* Falls through to STATE_DEEP_SLEEP conceptually; on many
         * platforms execution simply blocks in pmu_enter_deep_sleep()
         * until the RTC ISR sets g_rtc_wake_flag and wakes the core. */
        transition_to(STATE_DEEP_SLEEP);
        break;

    case STATE_DEEP_SLEEP:
        if (g_rtc_wake_flag) {
            g_rtc_wake_flag = false;
            g_ctx.wake_count++;
            transition_to(STATE_TELEMETRY);
        }
        break;

    case STATE_TELEMETRY:
        if (adc_sample_telemetry(&g_sample) == 0) {
            g_ctx.telemetry_write_pending = true;
        }

        if (g_ctx.telemetry_write_pending) {
            uint32_t slot = g_ctx.wake_count % EEPROM_LOG_SLOT_COUNT;
            if (eeprom_write_sample(&g_sample, slot) == 0) {
                g_ctx.telemetry_write_pending = false;
            } else {
                transition_to(STATE_ERROR);
                break;
            }
        }

        rtc_arm_next_wake(TELEMETRY_INTERVAL_SEC);
        transition_to(STATE_SLEEP_ENTER);
        break;

    case STATE_ACTIVE:
        /* Reserved for foreground/user-triggered active operation,
         * distinct from the periodic wake-sample-sleep cycle. */
        break;

    case STATE_ERROR:
        /* Fault handling: log, attempt recovery, or fall back to a
         * safe low-frequency retry rather than looping tightly. */
        rtc_arm_next_wake(TELEMETRY_INTERVAL_SEC * 4);
        transition_to(STATE_SLEEP_ENTER);
        break;
    }
}

int main(void)
{
    while (1) {
        state_machine_run();
    }
}
```

  

##   
Flash Health Monitor  
  

![](https://cdn.hashnode.com/uploads/covers/6560ea63862bec8ae76ef63f/c27718e7-0fab-406e-a3d3-1d1f1bd7c80f.png align="center")

![](https://cdn.hashnode.com/uploads/covers/6560ea63862bec8ae76ef63f/57135003-29e7-45a8-9877-c0f053d3aa53.png align="center")

```c
/*
 * flash_health_monitor.c
 *
 * TEMPLATE / REFERENCE CODE — Storage Infrastructure Flash Health &
 * Wear-Leveling Monitor
 *
 * Illustrates the structure described in the resume bullet:
 *   - Background flash-health monitoring against an MMC/eMMC device
 *   - EXT_CSD register parsing for wear-level indicators
 *   - RAM-buffered asynchronous write path to reduce flash wear
 *   - Threshold-based protection for degraded storage regions
 *
 * Generic, platform-agnostic skeleton meant to illustrate STRUCTURE and
 * control flow — not a drop-in service for a specific board or kernel
 * version. Adapt ioctl/sysfs calls to your actual MMC subsystem version.
 */

#include <stdint.h>
#include <stdbool.h>
#include <stddef.h>
#include <string.h>

/* ------------------------------------------------------------------------
 * 1. EXT_CSD parsing
 *
 * The eMMC spec (JEDEC JESD84) defines a 512-byte EXT_CSD register that the
 * device exposes over the MMC bus. Specific byte offsets report device
 * health — most notably byte 268 (DEVICE_LIFE_TIME_EST_TYP_A/B) and byte
 * 267 (PRE_EOL_INFO). These are read-only, standardized across vendors,
 * which is what makes EXT_CSD-based monitoring portable.
 * ---------------------------------------------------------------------- */

#define EXT_CSD_SIZE                 512
#define EXT_CSD_PRE_EOL_INFO         267
#define EXT_CSD_DEVICE_LIFE_TIME_A   268   /* SLC-type region wear estimate */
#define EXT_CSD_DEVICE_LIFE_TIME_B   269   /* MLC-type region wear estimate */

typedef enum {
    EOL_NORMAL        = 0x01,
    EOL_WARNING_80PCT = 0x02,   /* device has consumed ~80% of reserved
                                    spare blocks */
    EOL_URGENT_90PCT  = 0x03,   /* ~90%+ consumed — failure imminent */
} pre_eol_status_t;

typedef struct {
    pre_eol_status_t pre_eol;
    uint8_t          life_time_est_a;  /* 0x01 (0-10%) .. 0x0A (90-100%) */
    uint8_t          life_time_est_b;
} flash_health_t;

static int ext_csd_read_raw(const char *mmc_device_path, uint8_t *buf /* EXT_CSD_SIZE bytes */)
{
    /* On Linux this is typically done via mmc-utils (MMC_IOC_CMD ioctl,
     * CMD8 SEND_EXT_CSD) against /dev/mmcblkN, or by reading the kernel's
     * parsed sysfs nodes directly under
     * /sys/class/mmc_host/mmcX/mmcX:YYYY/ (life_time, pre_eol_info, etc.)
     * if the kernel driver already exposes them. Raw ioctl path shown
     * here for the case where direct register access is needed. */
    return mmc_send_ext_csd_ioctl(mmc_device_path, buf, EXT_CSD_SIZE);
}

static int flash_health_read(const char *mmc_device_path, flash_health_t *out)
{
    uint8_t ext_csd[EXT_CSD_SIZE];

    if (ext_csd_read_raw(mmc_device_path, ext_csd) != 0) {
        return -1;
    }

    out->pre_eol          = (pre_eol_status_t)ext_csd[EXT_CSD_PRE_EOL_INFO];
    out->life_time_est_a  = ext_csd[EXT_CSD_DEVICE_LIFE_TIME_A];
    out->life_time_est_b  = ext_csd[EXT_CSD_DEVICE_LIFE_TIME_B];
    return 0;
}

/* ------------------------------------------------------------------------
 * 2. Wear trend tracking
 *
 * A single EXT_CSD snapshot only tells you the current state. Tracking a
 * short history lets the monitor flag an abnormal wear *rate*, not just
 * an absolute threshold crossing — catching a runaway write pattern
 * before it reaches EOL_URGENT.
 * ---------------------------------------------------------------------- */

#define WEAR_HISTORY_DEPTH 16

typedef struct {
    flash_health_t samples[WEAR_HISTORY_DEPTH];
    uint32_t       sample_timestamps[WEAR_HISTORY_DEPTH];
    uint8_t        head;
    uint8_t        count;
} wear_history_t;

static wear_history_t g_wear_history = {0};

static void wear_history_push(const flash_health_t *sample, uint32_t timestamp)
{
    g_wear_history.samples[g_wear_history.head]           = *sample;
    g_wear_history.sample_timestamps[g_wear_history.head]  = timestamp;
    g_wear_history.head = (g_wear_history.head + 1) % WEAR_HISTORY_DEPTH;
    if (g_wear_history.count < WEAR_HISTORY_DEPTH) {
        g_wear_history.count++;
    }
}

/* Returns wear-level steps advanced per day, using the oldest and newest
 * samples in the ring. A sudden spike here indicates write amplification
 * or a runaway logging process rather than expected gradual wear. */
static float wear_history_rate_per_day(void)
{
    if (g_wear_history.count < 2) {
        return 0.0f;
    }
    uint8_t oldest_idx = (g_wear_history.head + WEAR_HISTORY_DEPTH - g_wear_history.count) % WEAR_HISTORY_DEPTH;
    uint8_t newest_idx = (g_wear_history.head + WEAR_HISTORY_DEPTH - 1) % WEAR_HISTORY_DEPTH;

    int32_t wear_delta = g_wear_history.samples[newest_idx].life_time_est_a
                        - g_wear_history.samples[oldest_idx].life_time_est_a;
    int32_t time_delta_sec = g_wear_history.sample_timestamps[newest_idx]
                            - g_wear_history.sample_timestamps[oldest_idx];

    if (time_delta_sec <= 0) {
        return 0.0f;
    }
    return (float)wear_delta / ((float)time_delta_sec / 86400.0f);
}

/* ------------------------------------------------------------------------
 * 3. RAM-buffered asynchronous write path
 *
 * Instead of committing every write immediately (one flash program
 * operation per write call), stage writes in a RAM buffer and flush in
 * batches. This trades a small, bounded data-loss-on-power-failure risk
 * for a large reduction in flash program/erase cycle count.
 * ---------------------------------------------------------------------- */

#define WRITE_BUFFER_CAPACITY_BYTES  (64 * 1024)
#define FLUSH_SIZE_THRESHOLD_BYTES   (32 * 1024)
#define FLUSH_TIME_THRESHOLD_SEC     30

typedef struct {
    uint8_t  data[WRITE_BUFFER_CAPACITY_BYTES];
    uint32_t used_bytes;
    uint32_t last_flush_timestamp;
} write_buffer_t;

static write_buffer_t g_write_buffer = {0};

typedef enum {
    REGION_HEALTHY = 0,
    REGION_DEGRADED,
    REGION_QUARANTINED,
} region_status_t;

/* Forward declaration — see section 4 */
static region_status_t region_status_for_offset(uint64_t offset);

static int flash_flush_buffer(const char *block_device_path)
{
    if (g_write_buffer.used_bytes == 0) {
        return 0;
    }

    /* Section 4 protection check happens before the physical write, not
     * after — the whole point is avoiding writes to a known-bad region. */
    uint64_t target_offset = flash_allocator_next_write_offset();
    if (region_status_for_offset(target_offset) == REGION_QUARANTINED) {
        target_offset = flash_allocator_next_safe_offset();  /* skip region */
    }

    int ret = block_device_write(block_device_path, target_offset,
                                  g_write_buffer.data, g_write_buffer.used_bytes);
    if (ret != 0) {
        return ret;   /* buffer retained; caller may retry */
    }

    g_write_buffer.used_bytes = 0;
    g_write_buffer.last_flush_timestamp = system_uptime_seconds();
    return 0;
}

/* Called on every incoming write request. Only touches flash once a
 * size or time threshold is crossed. */
static int flash_buffered_write(const char *block_device_path,
                                 const uint8_t *data, uint32_t len,
                                 uint32_t now)
{
    if (len > WRITE_BUFFER_CAPACITY_BYTES) {
        return -1;   /* oversized write bypasses buffering entirely */
    }

    if (g_write_buffer.used_bytes + len > WRITE_BUFFER_CAPACITY_BYTES) {
        int ret = flash_flush_buffer(block_device_path);
        if (ret != 0) return ret;
    }

    memcpy(&g_write_buffer.data[g_write_buffer.used_bytes], data, len);
    g_write_buffer.used_bytes += len;

    bool size_trip = g_write_buffer.used_bytes >= FLUSH_SIZE_THRESHOLD_BYTES;
    bool time_trip = (now - g_write_buffer.last_flush_timestamp) >= FLUSH_TIME_THRESHOLD_SEC;

    if (size_trip || time_trip) {
        return flash_flush_buffer(block_device_path);
    }
    return 0;
}

/* ------------------------------------------------------------------------
 * 4. Degraded-region protection
 *
 * Tracks regions that have crossed a wear or bad-block threshold and
 * steers future writes away from them, rather than waiting for the
 * filesystem layer to report an I/O error after damage already occurred.
 * ---------------------------------------------------------------------- */

#define MAX_TRACKED_REGIONS 64

typedef struct {
    uint64_t         start_offset;
    uint64_t         length;
    region_status_t  status;
    uint32_t         bad_block_count;
} storage_region_t;

static storage_region_t g_regions[MAX_TRACKED_REGIONS];
static uint32_t         g_region_count = 0;

#define BAD_BLOCK_DEGRADED_THRESHOLD    5
#define BAD_BLOCK_QUARANTINE_THRESHOLD  20

static void region_report_bad_block(uint64_t offset)
{
    for (uint32_t i = 0; i < g_region_count; i++) {
        storage_region_t *r = &g_regions[i];
        if (offset >= r->start_offset && offset < r->start_offset + r->length) {
            r->bad_block_count++;

            if (r->bad_block_count >= BAD_BLOCK_QUARANTINE_THRESHOLD) {
                r->status = REGION_QUARANTINED;
                log_warn("region at 0x%llx quarantined after %u bad blocks",
                         r->start_offset, r->bad_block_count);
            } else if (r->bad_block_count >= BAD_BLOCK_DEGRADED_THRESHOLD) {
                r->status = REGION_DEGRADED;
            }
            return;
        }
    }
}

static region_status_t region_status_for_offset(uint64_t offset)
{
    for (uint32_t i = 0; i < g_region_count; i++) {
        storage_region_t *r = &g_regions[i];
        if (offset >= r->start_offset && offset < r->start_offset + r->length) {
            return r->status;
        }
    }
    return REGION_HEALTHY;   /* untracked regions default to healthy */
}

/* ------------------------------------------------------------------------
 * 5. Background monitor loop
 * ---------------------------------------------------------------------- */

#define HEALTH_POLL_INTERVAL_SEC 3600   /* hourly EXT_CSD poll — frequent
                                            enough to catch trends, rare
                                            enough to add negligible load */

void flash_health_monitor_task(const char *mmc_device_path)
{
    flash_health_t health;
    uint32_t last_poll = 0;

    while (1) {
        uint32_t now = system_uptime_seconds();

        if (now - last_poll >= HEALTH_POLL_INTERVAL_SEC) {
            if (flash_health_read(mmc_device_path, &health) == 0) {
                wear_history_push(&health, now);

                if (health.pre_eol == EOL_URGENT_90PCT) {
                    log_error("flash device at 90%%+ EOL — escalating alert");
                    trigger_maintenance_alert(FLASH_ALERT_CRITICAL);
                } else if (health.pre_eol == EOL_WARNING_80PCT) {
                    log_warn("flash device at 80%%+ EOL — monitoring closely");
                    trigger_maintenance_alert(FLASH_ALERT_WARNING);
                }

                float rate = wear_history_rate_per_day();
                if (rate > EXPECTED_MAX_WEAR_RATE_PER_DAY) {
                    log_warn("abnormal wear rate detected: %.2f steps/day", rate);
                }
            }
            last_poll = now;
        }

        task_sleep_ms(1000);   /* background task tick — cooperative loop */
    }
}
```

  

## Multi-Stage Hardware Watchdog & Thermal Fail-Safe System  

![](https://cdn.hashnode.com/uploads/covers/6560ea63862bec8ae76ef63f/aac1b402-75a7-4a3f-8fda-395ce6e97d35.png align="center")

  

![](https://cdn.hashnode.com/uploads/covers/6560ea63862bec8ae76ef63f/7ff46dcf-bb78-434f-8ec8-c4c78cf41215.png align="center")

  

![](https://cdn.hashnode.com/uploads/covers/6560ea63862bec8ae76ef63f/6319cf03-53ea-46af-b3f5-f987aad06f0c.png align="center")

###   
  
watchdog\_thermal\_failsafe

```c

/*
 * watchdog_thermal_failsafe.c
 *
 * TEMPLATE / REFERENCE CODE — Multi-Stage Hardware Watchdog & Thermal
 * Fail-Safe System (KERNEL MODULE)
 *
 * Illustrates the structure described in the resume bullet:
 *   - Kernel-level hardware watchdog framework integration
 *   - Interrupt-driven watchdog kicking gated by a GPIO heartbeat signal
 *   - Thermal sensor threshold interrupt -> workqueue -> PMIC DVFS control
 *
 * "Multi-stage" here means layered fault detection:
 *   Stage 1 — userspace app proves liveness via a GPIO heartbeat pulse
 *   Stage 2 — kernel timer checks heartbeat freshness before kicking
 *             the hardware watchdog (soft layer, recoverable)
 *   Stage 3 — hardware watchdog timer itself, which forces an SoC reset
 *             if nobody kicks it (hard layer, last resort — fires even
 *             if the kernel itself is the thing that's hung)
 *
 * Generic, platform-agnostic skeleton meant to illustrate STRUCTURE and
 * control flow using standard Linux kernel framework APIs (watchdog core,
 * gpiod, threaded IRQs, workqueues, regulator/cpufreq for DVFS). Adapt
 * peripheral register offsets and platform-specific calls to your board.
 */

#include <linux/module.h>
#include <linux/platform_device.h>
#include <linux/watchdog.h>
#include <linux/gpio/consumer.h>
#include <linux/interrupt.h>
#include <linux/workqueue.h>
#include <linux/timer.h>
#include <linux/jiffies.h>
#include <linux/regulator/consumer.h>
#include <linux/cpufreq.h>
#include <linux/of.h>

#define DRIVER_NAME              "wdt_thermal_failsafe"
#define HEARTBEAT_STALE_JIFFIES  msecs_to_jiffies(2000)   /* stage-2 window:
                                     if no GPIO pulse in this window, stop
                                     kicking the hardware watchdog */
#define WATCHDOG_HW_TIMEOUT_SEC  10                        /* stage-3 window:
                                     hardware resets SoC if not kicked
                                     within this time regardless of kernel
                                     state */
#define THERMAL_TRIP_MC          85000   /* millidegrees C: DVFS trip point */
#define THERMAL_CRITICAL_MC      100000  /* millidegrees C: hard shutdown */

struct wdt_thermal_priv {
    struct watchdog_device wdd;
    struct gpio_desc      *heartbeat_gpio;
    int                    heartbeat_irq;
    unsigned long          last_heartbeat_jiffies;
    struct timer_list      supervise_timer;

    struct gpio_desc      *thermal_irq_gpio;   /* trip signal from sensor */
    int                    thermal_irq;
    struct work_struct     thermal_work;
    struct regulator      *pmic_core_rail;     /* DVFS voltage control */

    void __iomem          *wdt_regs;           /* platform HW watchdog MMIO */
};

/* ------------------------------------------------------------------------
 * Stage 1+2: GPIO heartbeat interrupt and supervising timer
 *
 * The heartbeat GPIO is driven by a userspace daemon (see companion
 * heartbeat_daemon.c) that only pulses it while its own liveness checks
 * pass. The kernel side never trusts a single pulse in isolation — it
 * tracks recency and only kicks the hardware watchdog if a pulse arrived
 * within the stale window, each time the supervising timer fires.
 * ---------------------------------------------------------------------- */

static irqreturn_t heartbeat_isr(int irq, void *data)
{
    struct wdt_thermal_priv *priv = data;

    /* Minimal ISR work: record the timestamp, defer everything else.
     * No I2C/blocking calls here — this runs in hard IRQ context. */
    priv->last_heartbeat_jiffies = jiffies;
    return IRQ_HANDLED;
}

static void wdt_hw_kick(struct wdt_thermal_priv *priv)
{
    /* Platform-specific: write the magic kick sequence/value to the
     * hardware watchdog's MMIO register to reset its countdown. */
    writel(WDT_KICK_MAGIC, priv->wdt_regs + WDT_KICK_REG_OFFSET);
}

static void supervise_timer_fn(struct timer_list *t)
{
    struct wdt_thermal_priv *priv = from_timer(priv, t, supervise_timer);
    unsigned long age = jiffies - priv->last_heartbeat_jiffies;

    if (age <= HEARTBEAT_STALE_JIFFIES) {
        wdt_hw_kick(priv);
    } else {
        /* Deliberately do NOT kick. The hardware watchdog's own
         * countdown (stage 3) will now continue toward reset unless a
         * fresh heartbeat arrives before WATCHDOG_HW_TIMEOUT_SEC elapses.
         * This is the layered part: a hung userspace app stops the
         * chain here even if the kernel itself is otherwise healthy. */
        pr_warn("%s: heartbeat stale (%ums old), withholding watchdog kick\n",
                DRIVER_NAME, jiffies_to_msecs(age));
    }

    mod_timer(&priv->supervise_timer, jiffies + msecs_to_jiffies(500));
}

/* ------------------------------------------------------------------------
 * Linux watchdog subsystem integration
 *
 * Registering against watchdog_device lets /dev/watchdogN exist and lets
 * userspace (or the kernel's own watchdog governor) interact with this
 * device using the standard WDIOC_* ioctl interface, instead of a bespoke
 * ABI.
 * ---------------------------------------------------------------------- */

static int wdt_start(struct watchdog_device *wdd)
{
    struct wdt_thermal_priv *priv = watchdog_get_drvdata(wdd);
    priv->last_heartbeat_jiffies = jiffies;
    wdt_hw_kick(priv);
    writel(WATCHDOG_HW_TIMEOUT_SEC, priv->wdt_regs + WDT_TIMEOUT_REG_OFFSET);
    writel(WDT_ENABLE_BIT, priv->wdt_regs + WDT_CTRL_REG_OFFSET);
    mod_timer(&priv->supervise_timer, jiffies + msecs_to_jiffies(500));
    return 0;
}

static int wdt_stop(struct watchdog_device *wdd)
{
    struct wdt_thermal_priv *priv = watchdog_get_drvdata(wdd);
    del_timer_sync(&priv->supervise_timer);
    writel(0, priv->wdt_regs + WDT_CTRL_REG_OFFSET);
    return 0;
}

static int wdt_ping(struct watchdog_device *wdd)
{
    /* Explicit userspace ping via WDIOC_KEEPALIVE ioctl on /dev/watchdogN.
     * Treated the same as a fresh GPIO heartbeat for supervising logic. */
    struct wdt_thermal_priv *priv = watchdog_get_drvdata(wdd);
    priv->last_heartbeat_jiffies = jiffies;
    return 0;
}

static const struct watchdog_ops wdt_ops = {
    .owner = THIS_MODULE,
    .start = wdt_start,
    .stop  = wdt_stop,
    .ping  = wdt_ping,
};

static const struct watchdog_info wdt_info = {
    .identity = DRIVER_NAME,
    .options  = WDIOF_KEEPALIVEPING | WDIOF_SETTIMEOUT,
};

/* ------------------------------------------------------------------------
 * Thermal trip interrupt -> workqueue -> PMIC DVFS
 *
 * The ISR only schedules work: I2C transactions to the PMIC and cpufreq
 * governor calls can sleep, so they cannot run in hard IRQ context.
 * ---------------------------------------------------------------------- */

static irqreturn_t thermal_trip_isr(int irq, void *data)
{
    struct wdt_thermal_priv *priv = data;
    schedule_work(&priv->thermal_work);
    return IRQ_HANDLED;
}

static void thermal_work_fn(struct work_struct *work)
{
    struct wdt_thermal_priv *priv = container_of(work, struct wdt_thermal_priv,
                                                   thermal_work);
    int temp_mc = thermal_sensor_read_temp_mc();   /* platform sensor read,
                                                        may be I2C/ADC based */

    if (temp_mc >= THERMAL_CRITICAL_MC) {
        pr_err("%s: critical temperature %d mC — forcing emergency shutdown\n",
               DRIVER_NAME, temp_mc);
        kernel_power_off();
        return;
    }

    if (temp_mc >= THERMAL_TRIP_MC) {
        /* Step down core voltage via the PMIC rail and drop CPU frequency
         * ceiling — DVFS response instead of a hard shutdown, so the
         * system keeps running in a degraded but functional state. */
        int uv = regulator_get_voltage(priv->pmic_core_rail);
        regulator_set_voltage(priv->pmic_core_rail, uv - DVFS_STEP_DOWN_UV,
                               uv - DVFS_STEP_DOWN_UV + DVFS_STEP_MARGIN_UV);
        cpufreq_update_policy(smp_processor_id());   /* re-clamp to a lower
                                                          max frequency set
                                                          via a registered
                                                          cpufreq notifier */
        pr_warn("%s: thermal trip at %d mC — DVFS throttle applied\n",
                DRIVER_NAME, temp_mc);
    }
}

/* ------------------------------------------------------------------------
 * Probe / remove
 * ---------------------------------------------------------------------- */

static int wdt_thermal_probe(struct platform_device *pdev)
{
    struct wdt_thermal_priv *priv;
    int ret;

    priv = devm_kzalloc(&pdev->dev, sizeof(*priv), GFP_KERNEL);
    if (!priv)
        return -ENOMEM;

    priv->wdt_regs = devm_platform_ioremap_resource(pdev, 0);
    if (IS_ERR(priv->wdt_regs))
        return PTR_ERR(priv->wdt_regs);

    priv->heartbeat_gpio = devm_gpiod_get(&pdev->dev, "heartbeat", GPIOD_IN);
    if (IS_ERR(priv->heartbeat_gpio))
        return PTR_ERR(priv->heartbeat_gpio);

    priv->heartbeat_irq = gpiod_to_irq(priv->heartbeat_gpio);
    ret = devm_request_irq(&pdev->dev, priv->heartbeat_irq, heartbeat_isr,
                            IRQF_TRIGGER_RISING, "wdt-heartbeat", priv);
    if (ret)
        return ret;

    priv->thermal_irq_gpio = devm_gpiod_get(&pdev->dev, "thermal-trip", GPIOD_IN);
    if (IS_ERR(priv->thermal_irq_gpio))
        return PTR_ERR(priv->thermal_irq_gpio);

    priv->thermal_irq = gpiod_to_irq(priv->thermal_irq_gpio);
    INIT_WORK(&priv->thermal_work, thermal_work_fn);
    ret = devm_request_threaded_irq(&pdev->dev, priv->thermal_irq,
                                     NULL, thermal_trip_isr,
                                     IRQF_TRIGGER_RISING | IRQF_ONESHOT,
                                     "wdt-thermal-trip", priv);
    if (ret)
        return ret;

    priv->pmic_core_rail = devm_regulator_get(&pdev->dev, "core");
    if (IS_ERR(priv->pmic_core_rail))
        return PTR_ERR(priv->pmic_core_rail);

    timer_setup(&priv->supervise_timer, supervise_timer_fn, 0);

    priv->wdd.info    = &wdt_info;
    priv->wdd.ops     = &wdt_ops;
    priv->wdd.timeout = WATCHDOG_HW_TIMEOUT_SEC;
    priv->wdd.parent  = &pdev->dev;
    watchdog_set_drvdata(&priv->wdd, priv);

    ret = devm_watchdog_register_device(&pdev->dev, &priv->wdd);
    if (ret)
        return ret;

    platform_set_drvdata(pdev, priv);
    pr_info("%s: probed, hw timeout %ds, heartbeat window %ums\n",
            DRIVER_NAME, WATCHDOG_HW_TIMEOUT_SEC,
            jiffies_to_msecs(HEARTBEAT_STALE_JIFFIES));
    return 0;
}

static int wdt_thermal_remove(struct platform_device *pdev)
{
    struct wdt_thermal_priv *priv = platform_get_drvdata(pdev);
    del_timer_sync(&priv->supervise_timer);
    cancel_work_sync(&priv->thermal_work);
    return 0;
}

static const struct of_device_id wdt_thermal_of_match[] = {
    { .compatible = "vendor,wdt-thermal-failsafe" },
    { }
};
MODULE_DEVICE_TABLE(of, wdt_thermal_of_match);

static struct platform_driver wdt_thermal_driver = {
    .probe  = wdt_thermal_probe,
    .remove = wdt_thermal_remove,
    .driver = {
        .name           = DRIVER_NAME,
        .of_match_table = wdt_thermal_of_match,
    },
};
module_platform_driver(wdt_thermal_driver);

MODULE_LICENSE("GPL");
MODULE_DESCRIPTION("Multi-stage hardware watchdog and thermal fail-safe driver");
```

###   
  
heartbeat\_daemon

```c
/*
 * heartbeat_daemon.c
 *
 * TEMPLATE / REFERENCE CODE — Multi-Stage Hardware Watchdog & Thermal
 * Fail-Safe System (USERSPACE COMPANION)
 *
 * This is Stage 1 of the layered design: a userspace daemon that only
 * pulses the heartbeat GPIO while its own application-level liveness
 * checks pass. The kernel module (watchdog_thermal_failsafe.c) never
 * trusts the GPIO line blindly — it only kicks the hardware watchdog
 * while pulses remain recent, so a hung userspace app (even if the
 * kernel itself is fine) still allows the hardware watchdog to fire.
 *
 * Also demonstrates the standard /dev/watchdog ioctl interface as an
 * alternative/complementary liveness channel to the GPIO path.
 *
 * Generic Linux userspace skeleton — adapt the health-check function and
 * GPIO sysfs paths (or libgpiod calls) to your actual platform.
 */

#include <stdio.h>
#include <stdlib.h>
#include <stdbool.h>
#include <unistd.h>
#include <fcntl.h>
#include <string.h>
#include <sys/ioctl.h>
#include <linux/watchdog.h>

#define HEARTBEAT_GPIO_SYSFS   "/sys/class/gpio/gpio60/value"  /* example
                                    export path — real deployments would
                                    more likely use libgpiod's
                                    gpiod_line_set_value() instead of raw
                                    sysfs, shown here for clarity */
#define WATCHDOG_DEV           "/dev/watchdog0"
#define HEARTBEAT_PERIOD_SEC   1

/* ------------------------------------------------------------------------
 * Application-level liveness checks
 *
 * The daemon only pulses the heartbeat if these checks pass — this is
 * what makes the heartbeat meaningful rather than a dumb periodic toggle
 * that would defeat the whole point of layered detection.
 * ---------------------------------------------------------------------- */

static bool critical_task_responsive(void)
{
    /* Example approaches, pick what fits the real system:
     *   - message queue round-trip to the critical task with a timeout
     *   - checking a shared-memory "last processed" timestamp the task
     *     updates itself, and comparing it against current time
     *   - a UNIX domain socket ping/pong to a supervised process
     * Kept abstract here since the concrete mechanism is app-specific. */
    return app_health_check_critical_task();
}

static bool storage_subsystem_healthy(void)
{
    /* Could tie into the flash health monitor's region tracker (see the
     * flash health monitor project) — refuse to heartbeat if storage
     * has gone into a state where continuing to run risks data loss. */
    return app_health_check_storage();
}

/* ------------------------------------------------------------------------
 * GPIO heartbeat pulse
 * ---------------------------------------------------------------------- */

static int gpio_pulse_heartbeat(void)
{
    int fd = open(HEARTBEAT_GPIO_SYSFS, O_WRONLY);
    if (fd < 0) {
        perror("open heartbeat gpio");
        return -1;
    }

    /* Rising edge — kernel ISR is configured for IRQF_TRIGGER_RISING,
     * so a 0 then 1 write sequence generates the interrupt the kernel
     * module is listening for. */
    if (write(fd, "0", 1) != 1) { close(fd); return -1; }
    usleep(1000);
    if (write(fd, "1", 1) != 1) { close(fd); return -1; }

    close(fd);
    return 0;
}

/* ------------------------------------------------------------------------
 * /dev/watchdog ioctl path (complementary/alternative channel)
 * ---------------------------------------------------------------------- */

static int watchdog_dev_ping(int wdt_fd)
{
    int dummy = 0;
    return ioctl(wdt_fd, WDIOC_KEEPALIVE, &dummy);
}

/* ------------------------------------------------------------------------
 * Main loop
 * ---------------------------------------------------------------------- */

int main(void)
{
    int wdt_fd = open(WATCHDOG_DEV, O_WRONLY);
    if (wdt_fd < 0) {
        perror("open /dev/watchdog0");
        /* Not necessarily fatal — GPIO path can operate independently,
         * but log loudly since losing this channel weakens the design. */
    }

    while (1) {
        bool healthy = critical_task_responsive() && storage_subsystem_healthy();

        if (healthy) {
            if (gpio_pulse_heartbeat() != 0) {
                fprintf(stderr, "heartbeat: GPIO pulse failed\n");
            }
            if (wdt_fd >= 0 && watchdog_dev_ping(wdt_fd) != 0) {
                perror("WDIOC_KEEPALIVE");
            }
        } else {
            /* Deliberately withhold both the GPIO pulse and the ioctl
             * ping. This is the intentional failure path — do nothing
             * and let the layered kernel/hardware timeout chain react,
             * rather than trying to handle the fault here in a process
             * that may itself be in a compromised state. */
            fprintf(stderr, "heartbeat: health check failed, withholding pulse\n");
        }

        sleep(HEARTBEAT_PERIOD_SEC);
    }

    if (wdt_fd >= 0) close(wdt_fd);
    return 0;
}
```
