>_0xFORUM
Sign in

Reproducing a file-watcher race with rr --chaos

in Debugging27 replies2.6k views

The bug only showed up on a loaded CI runner. Locally it was a heisenbug. rr record --chaos hit it in ~40 recordings.

Reverse-continue to the write, then watch the inode cache. The lost wakeup was in our own debounce, not the kernel. Happy to share the harness (MIT).

rr record --chaos ./watcher_test
rr replay -g watcher_test

Refs: rr

Lab / educational. Public binaries and patched classes only. Isolated VM.

// 27 REPLIES

This is the kind of thread that should be a sticky and is not. «Reproducing a file-watcher race with rr --chaos» — specifically The bug only showed up on a loaded CI runner. Dump the helper process. Always the helper process. I reproduced it on lab build 1297.

@blk

This is the kind of thread that should be a sticky and is not. «Reproducing a file-watcher race with rr --chaos» — specifically The bug only

I am not moving this to DMs so you can yell. Stay on the class. Not fully convinced yet. You wrote «The bug only showed up on a loaded CI runner». That is the sentence I keep. Dump the helper process. Always the helper process. I still have the snapshot named debu-53-pre.

Not fully convinced yet. You wrote «The bug only showed up on a loaded CI runner». That is the sentence I keep. !analyze is a hypothesis. !thread and the raw stacks are the evidence. My note id for this: 35-02.

@charly

Not fully convinced yet. You wrote «The bug only showed up on a loaded CI runner». That is the sentence I keep. !analyze is a hypothesis. !t

That is not what the listing shows. You are arguing a vibe. I dumped after OEP and then did this. «Reproducing a file-watcher race with rr --chaos» — specifically The bug only showed up on a loaded CI runner. rr --chaos is the first thing I try on a userspace race. If it cannot see it, I log TSC stamps. What did you key the join on — PID or process GUID? Version in my shot: current lab snapshot, not last year's blog.

This is the writeup I wanted when I was stuck. The load-bearing line: The bug only showed up on a loaded CI runner. Hang dump for hangs. Minidump for crashes I already understand. Hash of the public file, or are we arguing a shape? I still have the snapshot named debu-53-pre.

@cloudx

I dumped after OEP and then did this. «Reproducing a file-watcher race with rr --chaos» — specifically The bug only showed up on a loaded CI

I ran this on a licensed corpus binary. «Reproducing a file-watcher race with rr --chaos» — specifically The bug only showed up on a loaded CI runner. TTD queries that scan the whole trace are how you learn patience. Narrow the range. Same class as the January thread, different binary.

I want the listing, not the decompiler story. You wrote «The bug only showed up on a loaded CI runner». That is the sentence I keep. !analyze is a hypothesis. !thread and the raw stacks are the evidence. Same class as the March thread, different binary.

@ibrahimvibe

Did this on ARM64 last week — same shape, different pain. You wrote «The bug only showed up on a loaded CI runner». That is the sentence I k

Take the telegram pitch to the bin. Market listing or nothing. I reproduced it twice before I believed you. On «Reproducing a file-watcher race with rr --chaos»: The bug only showed up on a loaded CI runner. rr --chaos is the first thing I try on a userspace race. If it cannot see it, I log TSC stamps. Can you quote the offset instead of the graph screenshot? Same class as the March thread, different binary.

@loadx

I will argue the opposite and then probably agree. The load-bearing line: The bug only showed up on a loaded CI runner. !analyze is a hypoth

You are treating a checksum as a signature again. Same wall I hit last quarter. «Reproducing a file-watcher race with rr --chaos» — specifically The bug only showed up on a loaded CI runner. Kernel time travel is not user TTD. Stepping into a syscall will not take you to the kernel. I wrote a 12-line script and then threw it away. The listing was enough.

@flarez

This is the writeup I wanted when I was stuck. The load-bearing line: The bug only showed up on a loaded CI runner. Hang dump for hangs. Min

Do not call people skids because they use Ghidra. This belongs in the first-hour ritual. «Reproducing a file-watcher race with rr --chaos» — specifically The bug only showed up on a loaded CI runner. !analyze is a hypothesis. !thread and the raw stacks are the evidence. Same class as the March thread, different binary.

@cryptb

I ran this on a licensed corpus binary. «Reproducing a file-watcher race with rr --chaos» — specifically The bug only showed up on a loaded

I read the patch. You read a tweet. Those are not the same source. I ran this on a licensed corpus binary. The load-bearing line: The bug only showed up on a loaded CI runner. TTD queries that scan the whole trace are how you learn patience. Narrow the range. If anyone DMs me a zip I will not open it. Hash in-thread.

@grid

The screenshot is the useful part of the post. The load-bearing line: The bug only showed up on a loaded CI runner. TTD queries that scan th

You are describing a live target. Stop. Patched class only. If you only have the decompiler, you do not have the bug. The load-bearing line: The bug only showed up on a loaded CI runner. !analyze is a hypothesis. !thread and the raw stacks are the evidence. Version in my shot: current lab snapshot, not last year's blog.

@freshmode

This belongs in the first-hour ritual. «Reproducing a file-watcher race with rr --chaos» — specifically The bug only showed up on a loaded C

The screenshot is the useful part of the post. The load-bearing line: The bug only showed up on a loaded CI runner. TTD queries that scan the whole trace are how you learn patience. Narrow the range. Same class as the October thread, different binary.

@dfir

I want the listing, not the decompiler story. You wrote «The bug only showed up on a loaded CI runner». That is the sentence I keep. !analyz

Call-convention guess is not evidence. The screenshot is the useful part of the post. The load-bearing line: The bug only showed up on a loaded CI runner. SetThreadDescription is free. I will keep nagging. Version in my shot: current lab snapshot, not last year's blog.

I still keep a paper notebook for this kind of note. You wrote «The bug only showed up on a loaded CI runner». That is the sentence I keep. TTD queries that scan the whole trace are how you learn patience. Narrow the range. I wrote a 12-line script and then threw it away. The listing was enough.

Did this on ARM64 last week — same shape, different pain. You wrote «The bug only showed up on a loaded CI runner». That is the sentence I keep. Dump the helper process. Always the helper process. I wrote a 12-line script and then threw it away. The listing was enough.

Also: SetThreadDescription is free. I will keep nagging.

@labsec

This is the kind of thread that should be a sticky and is not. «Reproducing a file-watcher race with rr --chaos» — specifically The bug only

I will argue the opposite and then probably agree. The load-bearing line: The bug only showed up on a loaded CI runner. !analyze is a hypothesis. !thread and the raw stacks are the evidence. Pinned a comment at 0x1400047df in the listing.

@kelvinpro

I still keep a paper notebook for this kind of note. You wrote «The bug only showed up on a loaded CI runner». That is the sentence I keep.

Quote the bytes or sit down. This is the kind of thread that should be a sticky and is not. «Reproducing a file-watcher race with rr --chaos» — specifically The bug only showed up on a loaded CI runner. TTD queries that scan the whole trace are how you learn patience. Narrow the range. Same class as the January thread, different binary.

Did this on ARM64 last week — same shape, different pain. The load-bearing line: The bug only showed up on a loaded CI runner. !analyze is a hypothesis. !thread and the raw stacks are the evidence. Can you quote the offset instead of the graph screenshot? Version in my shot: current lab snapshot, not last year's blog.

@nexor

Did this on ARM64 last week — same shape, different pain. The load-bearing line: The bug only showed up on a loaded CI runner. !analyze is a

You skipped isolation and then asked why the box is dirty. That is on you. Please keep the hashes and drop the mystery zips. On «Reproducing a file-watcher race with rr --chaos»: The bug only showed up on a loaded CI runner. !analyze is a hypothesis. !thread and the raw stacks are the evidence. Pinned a comment at 0x14000670d in the listing.

This matches a public n-day class from last patch Tuesday. You wrote «The bug only showed up on a loaded CI runner». That is the sentence I keep. Page heap and ASan catch different lies. I run both. I will +rep a listing and −rep a vibe. That is the deal.

Also: Hang dump for hangs. Minidump for crashes I already understand.

@paulflex

This matches a public n-day class from last patch Tuesday. You wrote «The bug only showed up on a loaded CI runner». That is the sentence I

I am not moving this to DMs so you can yell. Stay on the class. I would have written the opposite conclusion a year ago. The load-bearing line: The bug only showed up on a loaded CI runner. rr --chaos is the first thing I try on a userspace race. If it cannot see it, I log TSC stamps. I wrote a 12-line script and then threw it away. The listing was enough.

@pwned

I would have written the opposite conclusion a year ago. The load-bearing line: The bug only showed up on a loaded CI runner. rr --chaos is

Agreed on the class, not on the tool. You wrote «The bug only showed up on a loaded CI runner». That is the sentence I keep. SetThreadDescription is free. I will keep nagging. My note id for this: 35-22.

@riskr

Agreed on the class, not on the tool. You wrote «The bug only showed up on a loaded CI runner». That is the sentence I keep. SetThreadDescri

That is not what the listing shows. You are arguing a vibe. Quietly the best note on this board this month. The load-bearing line: The bug only showed up on a loaded CI runner. Dump the helper process. Always the helper process. Did you snapshot before, or is this a restore-from-memory story? My note id for this: 35-23.

Good. Dated shot, version in the post. «Reproducing a file-watcher race with rr --chaos» — specifically The bug only showed up on a loaded CI runner. TTD queries that scan the whole trace are how you learn patience. Narrow the range. I will +rep a listing and −rep a vibe. That is the deal.

@sn0op

Good. Dated shot, version in the post. «Reproducing a file-watcher race with rr --chaos» — specifically The bug only showed up on a loaded C

I read the patch. You read a tweet. Those are not the same source. Good. Dated shot, version in the post. You wrote «The bug only showed up on a loaded CI runner». That is the sentence I keep. Page heap and ASan catch different lies. I run both. If anyone DMs me a zip I will not open it. Hash in-thread.

Not fully convinced yet. «Reproducing a file-watcher race with rr --chaos» — specifically The bug only showed up on a loaded CI runner. Hang dump for hangs. Minidump for crashes I already understand. I will +rep a listing and −rep a vibe. That is the deal.

Sign in to reply. Guests can read reversing, pentesting, coding and greyhat threads.