What happens when ptrace SINGLESTEP is called in aarch64, Linux kernel?
Linux reference for this question: 5.15.5 (latest stable in November 2021).
Nomenclature: tracer (the process which traces) and tracee (the traced process).
With static analysis + ftracing the linux kernel, I tried to reconstruct what precisely happens upon a ptrace SINGLESTEP call. I tried to write what I understood in a learning OS, failing in obtaining the same behavious. Before asking my question, let me summarize the process that I tried to reconstruct:
- Single stepping is enabled in debug_monitor.c:
/* ptrace API */
void user_enable_single_step(struct task_struct *task)
{
struct thread_info *ti = task_thread_info(task);
if (!test_and_set_ti_thread_flag(ti, TIF_SINGLESTEP))
set_regs_spsr_ss(task_pt_regs(task));
}
- This, in arm64, consists in setting the single step bit SPSR_EL1.SS (position 21):
/*
* Single step API and exception handling.
*/
static void set_user_regs_spsr_ss(struct user_pt_regs *regs)
{
regs->pstate |= DBG_SPSR_SS;
}
- This, I imagine, should raise a "debug exception" (user level, EL0), catched and handled in entry-common.c:
static void noinstr el0_dbg(struct pt_regs *regs, unsigned long esr)
{
/* Only watchpoints write FAR_EL1, otherwise its UNKNOWN */
unsigned long far = read_sysreg(far_el1);
enter_from_user_mode(regs);
do_debug_exception(far, esr, regs);
local_daif_restore(DAIF_PROCCTX);
exit_to_user_mode(regs);
}
do_debug_exception()is defined in fault.c and, since it is a software step, it should call the functionearly_brk64of debug_fault_info data structure:
/*
* __refdata because early_brk64 is __init, but the reference to it is
* clobbered at arch_initcall time.
* See traps.c and debug-monitors.c:debug_traps_init().
*/
static struct fault_info __refdata debug_fault_info[] = {
{ do_bad, SIGTRAP, TRAP_HWBKPT, "hardware breakpoint" },
{ do_bad, SIGTRAP, TRAP_HWBKPT, "hardware single-step" },
{ do_bad, SIGTRAP, TRAP_HWBKPT, "hardware watchpoint" },
{ do_bad, SIGKILL, SI_KERNEL, "unknown 3" },
{ do_bad, SIGTRAP, TRAP_BRKPT, "aarch32 BKPT" },
{ do_bad, SIGKILL, SI_KERNEL, "aarch32 vector catch" },
{ early_brk64, SIGTRAP, TRAP_BRKPT, "aarch64 BRK" },
{ do_bad, SIGKILL, SI_KERNEL, "unknown 7" },
};
- The latter is defined in traps.c, it calls bug_handler function, which in turn calls arm64_skip_faulting_instruction(): the latter updates the
PC(of the tracee) of 4 bytes (aarch64 instructions are on 32 bits):
void arm64_skip_faulting_instruction(struct pt_regs *regs, unsigned long size)
{
regs->pc += size;
/*
* If we were single stepping, we want to get the step exception after
* we return from the trap.
*/
if (user_mode(regs))
user_fastforward_single_step(current);
if (compat_user_mode(regs))
advance_itstate(regs);
else
regs->pstate &= ~PSR_BTYPE_MASK;
}
- Finally, user_fastforward_single_step is called, which in turns calls clear_user_regs_spsr_ss which just reset the SS bit previously set:
static void clear_user_regs_spsr_ss(struct user_pt_regs *regs)
{
regs->pstate &= ~DBG_SPSR_SS;
}
If this chain of calls is correct, I can't understand where and how the context switching (from tracer to tracee) happens. Indeed, from these calls, it seems that points 2-6 regards the tracer. I can't notice a context switch in this chain, but it should.
I have tried to replicate all these steps in a learning OS. For the sake of clarity, I failed generating an exception after setting SPSR.SS bit (point 2), but I forced to generate an hardware debug exception by setting SS bit, MDE bit and KDE bit on MDSCR_EL1 register.
Once I did this trick, my steps were verified but, indeed, the exception was captured by the tracer (and not by the tracee as it should): PC of the tracer were updated, and so on. I think that steps that I retrieved through static analysis of the code and ftrace are not completely correct. Can you help in identifying where?