<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://goingforstudying-ctrl.github.io/feed.xml" rel="self" type="application/atom+xml" /><link href="https://goingforstudying-ctrl.github.io/" rel="alternate" type="text/html" /><updated>2026-10-11T18:14:55+00:00</updated><id>https://goingforstudying-ctrl.github.io/feed.xml</id><title type="html">Field Notes</title><subtitle>Machine learning and AI engineering notes: gradients, retrieval pipelines, inference systems, and distributed backends.</subtitle><author><name>goingforstudying-ctrl</name></author><entry><title type="html">A late success sent Airflow down the failure path</title><link href="https://goingforstudying-ctrl.github.io/notes/airflow-stale-success/" rel="alternate" type="text/html" title="A late success sent Airflow down the failure path" /><published>2026-10-11T00:00:00+00:00</published><updated>2026-10-11T00:00:00+00:00</updated><id>https://goingforstudying-ctrl.github.io/notes/airflow-stale-success</id><content type="html" xml:base="https://goingforstudying-ctrl.github.io/notes/airflow-stale-success/"><![CDATA[<p>The worker reported <code class="language-plaintext highlighter-rouge">SUCCESS</code>. Airflow could respond by taking the task down its failure path.</p>

<p>The report was accurate: the worker had exited after deferring the task. It was also late. Before the scheduler processed it, the trigger had resumed the task and its continuation had reached <code class="language-plaintext highlighter-rouge">QUEUED</code>. The event and the database state described different moments in the same execution.</p>

<p>What bothered me was that neither side had to be lying. The scheduler had enough information to recognize the ordering, but its existing guard stopped one state too early.</p>

<h2 id="two-timelines-meet-at-the-scheduler">Two timelines meet at the scheduler</h2>

<p>A deferrable task releases its worker while waiting for a trigger. Resuming the task and processing the worker’s executor event can progress independently. That permits this ordering:</p>

<figure class="technical-figure" aria-labelledby="defer-figure-caption">
  <div class="technical-figure-heading"><span class="eyebrow">Fig. 01 / Airflow</span><span class="figure-tag">A late event</span></div>
  <div class="event-timeline">
    <div class="timeline-step"><span class="timeline-number" aria-hidden="true">01</span><div><span class="diagram-kicker">Worker</span><strong>Defers and exits</strong><p>Emits <code>SUCCESS</code>; the scheduler hasn't processed it.</p></div><span class="state-badge">DEFERRED</span></div>
    <div class="timeline-step"><span class="timeline-number" aria-hidden="true">02</span><div><span class="diagram-kicker">Trigger, then scheduler</span><strong>Resumes the continuation</strong><p>The task advances through <code>SCHEDULED</code>.</p></div><span class="state-badge">QUEUED</span></div>
    <div class="timeline-step"><span class="timeline-number" aria-hidden="true">03</span><div><span class="diagram-kicker">Scheduler</span><strong>Reads the older SUCCESS</strong><p><code>next_method</code> still identifies the continuation.</p></div><span class="state-badge">QUEUED</span></div>
  </div>
  <div class="timeline-verdict"><span class="diagram-kicker">The patched guard</span><strong>Ignore this event; keep the continuation queued.</strong></div>
  <figcaption id="defer-figure-caption">The executor event belongs to the worker's defer exit. The current state belongs to the resumed task.</figcaption>
</figure>

<p>By the time the older <code class="language-plaintext highlighter-rouge">SUCCESS</code> arrives, <code class="language-plaintext highlighter-rouge">QUEUED</code> describes the continuation. Treating that pair as an ordinary mismatch can punish a task that has already made progress.</p>

<p>For an ML workflow waiting on an external job or an artifact, deferral is a useful way to avoid occupying a worker during the wait. The same lifecycle gives the control plane more than one event to reconcile. A success from one phase doesn’t tell the scheduler what the resumed phase should do next.</p>

<p>I find it easier to reason about this race with the events in order than with a list of state names. <code class="language-plaintext highlighter-rouge">SUCCESS</code> and <code class="language-plaintext highlighter-rouge">QUEUED</code> sound contradictory in isolation. Place the defer exit before the trigger, and the contradiction disappears.</p>

<h2 id="the-continuation-marker-narrows-the-exception">The continuation marker narrows the exception</h2>

<p>Airflow already handled the resumed task while it was <code class="language-plaintext highlighter-rouge">SCHEDULED</code>. The existing check also required <code class="language-plaintext highlighter-rouge">next_method</code>, which identifies the continuation, before ignoring the stale executor success. The missing case was a continuation that had advanced to <code class="language-plaintext highlighter-rouge">QUEUED</code>.</p>

<p>I extended the state check to include both. Written out, the resume-after-defer exception requires:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="p">(</span>
    <span class="n">ti</span><span class="p">.</span><span class="n">state</span> <span class="ow">in</span> <span class="p">(</span><span class="n">TaskInstanceState</span><span class="p">.</span><span class="n">SCHEDULED</span><span class="p">,</span> <span class="n">TaskInstanceState</span><span class="p">.</span><span class="n">QUEUED</span><span class="p">)</span>
    <span class="ow">and</span> <span class="n">state</span> <span class="o">==</span> <span class="n">TaskInstanceState</span><span class="p">.</span><span class="n">SUCCESS</span>
    <span class="ow">and</span> <span class="n">ti</span><span class="p">.</span><span class="n">next_method</span> <span class="ow">is</span> <span class="ow">not</span> <span class="bp">None</span>
<span class="p">)</span>
</code></pre></div></div>

<p>Each condition does work. The current state must belong to the resumed path. The arriving event must be a success. The task must still carry a continuation marker. A queued task without that marker can represent a real mismatch and should retain the existing handling.</p>

<p>Keeping those conditions together mattered more to me than making the check shorter. This patch recognizes a specific ordering; it doesn’t change the scheduler’s general treatment of state disagreements.</p>

<h2 id="remove-the-reason-for-the-exception">Remove the reason for the exception</h2>

<p>The regression sets up a queued task with <code class="language-plaintext highlighter-rouge">next_method = "execute_callback"</code>, then delivers the older success event. The task must remain queued. The scheduler must send no failure callback and record no unexpected metric.</p>

<p>That checks the race, but a guard that ignored every success for every queued task could pass it too. The second half removes <code class="language-plaintext highlighter-rouge">next_method</code> and delivers the same event again. This time the scheduler must record <code class="language-plaintext highlighter-rouge">scheduler.tasks.killed_externally</code>.</p>

<p>That is the assertion I would look for first in review. The test changes only the evidence that justified the exception, then asks the scheduler to resume its ordinary behavior. It makes an overly broad fix visible without needing a second, unrelated setup.</p>

<p>The implementation is a small extension to a state check. Its scope is easier to trust because the regression tests both the continuation and the case where the continuation marker is absent.</p>

<p>When I read orchestration code now, I ask which phase an event belongs to before deciding whether its state conflicts with the database. A pipeline can spend much longer waiting for external work than running Python. Correctly handing that work back to the scheduler is part of keeping the pipeline reliable.</p>

<p>My <a href="https://github.com/apache/airflow/pull/68741">Apache Airflow PR #68741</a> merged on July 2, 2026. <a href="https://github.com/apache/airflow/pull/68741/files">State guard and regression test</a>.</p>]]></content><author><name>goingforstudying-ctrl</name></author><category term="distributed-systems" /><category term="state-machines" /><summary type="html"><![CDATA[The worker's success report was accurate. By the time the scheduler read it, the task had already resumed, and the state guard was one state short.]]></summary></entry><entry><title type="html">The socket race started in a destructor</title><link href="https://goingforstudying-ctrl.github.io/notes/celery-pubsub-concurrency/" rel="alternate" type="text/html" title="The socket race started in a destructor" /><published>2026-10-11T00:00:00+00:00</published><updated>2026-10-11T00:00:00+00:00</updated><id>https://goingforstudying-ctrl.github.io/notes/celery-pubsub-concurrency</id><content type="html" xml:base="https://goingforstudying-ctrl.github.io/notes/celery-pubsub-concurrency/"><![CDATA[<p>The task could finish, and collecting its result could still lose a fight over the Redis socket. One caller was Celery’s drainer, waiting for an update. The other was an <code class="language-plaintext highlighter-rouge">AsyncResult</code> being destroyed.</p>

<p>Under gevent, overlapping access to the shared pub/sub object could raise <code class="language-plaintext highlighter-rouge">ConcurrentObjectUseError</code>. Following the second caller took me out of the polling loop and into <code class="language-plaintext highlighter-rouge">AsyncResult.__del__</code>, through result removal, and finally to <code class="language-plaintext highlighter-rouge">unsubscribe()</code>.</p>

<p>I had been reading cleanup as code that ran after the interesting work. Here it was another active participant in the protocol, issuing I/O while the drainer was still using the connection. That changed where I put the lock, and which kind of lock the consumer needed.</p>

<h2 id="cleanup-is-a-socket-user">Cleanup is a socket user</h2>

<p>The drainer calls <code class="language-plaintext highlighter-rouge">get_message()</code> to collect task updates. Result destruction can reach <code class="language-plaintext highlighter-rouge">remove_pending_result</code>, then <code class="language-plaintext highlighter-rouge">cancel_for</code>, and eventually <code class="language-plaintext highlighter-rouge">pubsub.unsubscribe()</code>. Redis-py pub/sub objects don’t support concurrent use, so protecting only the polling call leaves the other caller free to enter the same object.</p>

<p>The patch gives <code class="language-plaintext highlighter-rouge">ResultConsumer</code> a shared <code class="language-plaintext highlighter-rouge">RLock</code> around pub/sub access, including subscription changes, polling, reconnection, and closing the object. The subscription set is updated under that lock too. The ownership rule needs to cover the consumer’s view of the subscriptions as well as the socket calls.</p>

<figure class="technical-figure" aria-labelledby="pubsub-figure-caption">
  <div class="technical-figure-heading"><span class="eyebrow">Fig. 01 / Result consumer</span><span class="figure-tag">After the patch</span></div>
  <div class="diagram-branches">
    <div class="diagram-card"><span class="diagram-kicker">Drainer</span><strong>Poll for a result</strong><code>get_message()</code></div>
    <div class="diagram-card"><span class="diagram-kicker">Result cleanup</span><strong>Remove a subscription</strong><code>unsubscribe()</code></div>
  </div>
  <div class="ownership-gate"><span class="diagram-arrow" aria-hidden="true">↓</span><span>One shared <code>RLock</code></span><span class="diagram-arrow" aria-hidden="true">↓</span></div>
  <div class="diagram-card success-card socket-card"><span class="diagram-kicker">Shared Redis pub/sub object</span><strong>One caller owns the socket at a time</strong></div>
  <div class="nested-callback"><span class="diagram-kicker">Reentry while processing a message</span><div><code>on_state_change</code><span aria-hidden="true">→</span><code>cancel_for</code><span aria-hidden="true">→</span><code>unsubscribe</code></div><p>The nested call acquires the same lock again.</p></div>
  <figcaption id="pubsub-figure-caption">The lock covers both callers. Reentrancy lets a callback unsubscribe while the outer drainer call still holds it.</figcaption>
</figure>

<p>I find the destructor path useful to keep visible in a diagram. A reviewer looking only for public methods that explicitly subscribe or poll can miss a caller whose entry point is object lifetime.</p>

<h2 id="the-callback-comes-back-through-the-same-lock">The callback comes back through the same lock</h2>

<p>While holding the lock, the drainer can process a message through <code class="language-plaintext highlighter-rouge">on_state_change</code>. If the task is ready, that callback can call <code class="language-plaintext highlighter-rouge">cancel_for</code>, which needs to unsubscribe. The consumer reaches the lock again before the outer call has released it.</p>

<p>A plain mutex would make the consumer wait on itself. The reentrant lock lets that nested call proceed under the same ownership. Reconnection also reaches nested consumer operations, so choosing <code class="language-plaintext highlighter-rouge">RLock</code> requires following the callbacks, not just counting concurrent callers.</p>

<p>That was the most revealing part of the review for me. The declaration fits on one line; the reason for it sits several calls away. Adding synchronization without tracing those calls would trade a socket race for a deadlock.</p>

<p>Two less visible boundaries matter too. When no pub/sub object exists, <code class="language-plaintext highlighter-rouge">drain_events</code> sleeps outside the lock. Otherwise, a caller trying to start consumption would have to wait through an idle sleep. After a fork, <code class="language-plaintext highlighter-rouge">on_after_fork</code> creates a fresh lock before cleanup: the inherited lock may have been held by a thread that doesn’t exist in the child.</p>

<h2 id="test-ownership-before-testing-the-whole-queue">Test ownership before testing the whole queue</h2>

<p>The unit test uses a fake pub/sub object that records overlapping calls. A drainer polls while other threads subscribe and unsubscribe; the fake makes concurrent entry observable. That gives the regression a precise failure condition without relying on a real socket race to happen at the right moment.</p>

<p>Other unit checks cover lock ownership on the relevant paths, sleeping outside the lock when idle, and replacing the inherited lock after a fork. The PR also adds Redis integration tests for concurrent result collection and subscription churn. Those exercise the connection behavior the fake can’t reproduce.</p>

<p>I want both layers. The fake helps explain what broke when a test fails. The integration test checks that the ownership rule still makes sense against Redis, including callers entering through the normal result API.</p>

<h2 id="serialization-has-a-latency-cost">Serialization has a latency cost</h2>

<p>Polling holds the lock while waiting for a message. A new subscription or cancellation must wait for that poll to return. The PR discusses a drainer poll timeout of up to one second; that timeout describes the poll, not an end-to-end subscription latency guarantee.</p>

<p>For a queue feeding embedding jobs or offline model evaluations, I would measure result collection and subscription latency alongside worker execution time. Even after a fast worker finishes, the caller can still be waiting behind a poll in the result consumer.</p>

<p>A design in which the drainer owns the socket and accepts subscription commands through a queue could move that boundary. This patch keeps the existing consumer structure and serializes its callers. It fixes ownership at a smaller scope, with a waiting cost that a deployment can measure.</p>

<p>My <a href="https://github.com/celery/celery/pull/10671">Celery PR #10671</a> merged on September 21, 2026. <a href="https://github.com/celery/celery/pull/10671/files">Implementation and tests</a>.</p>]]></content><author><name>goingforstudying-ctrl</name></author><category term="concurrency" /><category term="task-queues" /><summary type="html"><![CDATA[I followed a Redis socket race out of Celery's polling loop and into result cleanup. The fix needed a lock that could survive callbacks and a fork.]]></summary></entry><entry><title type="html">The retry skipped the document that failed</title><link href="https://goingforstudying-ctrl.github.io/notes/swiftide-indexing-cache/" rel="alternate" type="text/html" title="The retry skipped the document that failed" /><published>2026-10-11T00:00:00+00:00</published><updated>2026-10-11T00:00:00+00:00</updated><id>https://goingforstudying-ctrl.github.io/notes/swiftide-indexing-cache</id><content type="html" xml:base="https://goingforstudying-ctrl.github.io/notes/swiftide-indexing-cache/"><![CDATA[<p>The second run was supposed to recover the missing document. Instead, it skipped it.</p>

<p>An embedding or storage failure should leave work for a retry. In Swiftide, <code class="language-plaintext highlighter-rouge">filter_cached</code> could write the cache entry before either stage had finished. The next run saw that entry and treated the document as handled. A temporary downstream failure had become a persistent reason to leave the document out of the index.</p>

<p>My first sketch was embarrassingly short: move the cache write to the end. Then I tried to define what “the end” meant for a document that became several chunks, lost an error to <code class="language-plaintext highlighter-rouge">filter_errors</code>, and passed through a merged pipeline. Most of the work in this patch lives in that definition.</p>

<h2 id="the-cache-was-claiming-success-too-early">The cache was claiming success too early</h2>

<p>The original sequence was a cache miss, an immediate cache write, then downstream processing. If embedding or storage failed, the cache still remembered the early write. A retry faithfully skipped the source that needed another attempt.</p>

<p>For a retrieval pipeline, that failure is uncomfortable because the index can remain usable. Other documents are searchable, and a model can still produce an answer. The missing source shows up later as missing evidence. When debugging that kind of gap, I’d check whether the source reached the index before spending time on the embedding model or prompt.</p>

<p>The patch makes the filter record which sources need work. Cache writes happen in <code class="language-plaintext highlighter-rouge">run()</code>, after the stream finishes without an unhandled error. Nodes reaching the end provide completion candidates; the final decision also checks failures and which cache filters each source actually passed.</p>

<h2 id="chunking-changes-the-unit-of-completion">Chunking changes the unit of completion</h2>

<p>A cache lookup starts with a source document. A chunker can turn it into several derived nodes. The storage backend sees those children, so the pipeline needs to carry the original source identity all the way through.</p>

<p>Derived nodes now preserve the first source ancestor in <code class="language-plaintext highlighter-rouge">parent_id</code>. Redis, redb, and DuckDB cache integrations use that identity, and completion candidates are deduplicated. Several successful children can lead to one cache write for their source.</p>

<p>One successful child still doesn’t prove the source finished. This is the regression I find most useful:</p>

<figure class="technical-figure" aria-labelledby="indexing-figure-caption">
  <div class="technical-figure-heading"><span class="eyebrow">Fig. 01 / Retrieval pipeline</span><span class="figure-tag">Partial success</span></div>
  <div class="completion-flow">
    <div class="diagram-card source-card"><span class="diagram-kicker">Source A</span><strong>One document</strong><span>Passes the cache filter</span></div>
    <span class="diagram-arrow" aria-hidden="true">↓</span>
    <div class="diagram-branches">
      <div class="diagram-card success-card"><span class="diagram-kicker">Chunker yields a child</span><strong>Child is stored</strong><span>Completion candidate: A</span></div>
      <div class="diagram-card warning-card"><span class="diagram-kicker">Chunker then errors</span><strong>Source A is recorded failed</strong><span><code>filter_errors</code> drops the error</span></div>
    </div>
    <span class="diagram-arrow" aria-hidden="true">↓</span>
    <div class="diagram-verdict"><span><code>run()</code> returns <code>Ok</code></span><strong>Source A stays uncached</strong><span>The next run can retry the source.</span></div>
  </div>
  <figcaption id="indexing-figure-caption">A successful child and a clean return can coexist with a failed source. The failure registry keeps that source eligible for retry.</figcaption>
</figure>

<p>The chunker yields a child successfully, then returns an error. With <code class="language-plaintext highlighter-rouge">filter_errors</code>, the pipeline can store that child and return successfully. If completion relied only on the returned status and the stored child, the parent would be cached even though its fan-out failed partway through.</p>

<p>The patch records failed source IDs before the error can disappear from the stream. That history survives <code class="language-plaintext highlighter-rouge">filter_errors</code>, and the final marking phase excludes the source. The successful child stays stored; the source stays eligible for retry.</p>

<p>I like this test because it makes two plausible success signals insufficient on purpose. A stored node and an <code class="language-plaintext highlighter-rouge">Ok</code> result are both real. Neither one says that every part of the source completed.</p>

<h2 id="a-merged-pipeline-needs-branch-history">A merged pipeline needs branch history</h2>

<p>There is another way to overstate completion. Suppose a source passes through the right branch’s cache filter, then the branches merge. Marking it in the left branch’s cache would claim work that branch never performed. A later change to the split predicate could then skip the source on the left.</p>

<p>The completion bookkeeping therefore uses pairs of cache registration and source ID. A completed source only updates a cache whose filter it actually passed. Independently built pipelines also link their failure and cache-pass registries when they merge, so the final decision can see both sides’ history.</p>

<p>The rule is easier to review when I write out its conditions:</p>

<table>
  <thead>
    <tr>
      <th style="text-align: left">Evidence at the end of <code class="language-plaintext highlighter-rouge">run()</code></th>
      <th style="text-align: left">What it permits</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td style="text-align: left">The stream ended without an unhandled error</td>
      <td style="text-align: left">Enter the cache-marking phase</td>
    </tr>
    <tr>
      <td style="text-align: left">A source has a descendant that reached the end</td>
      <td style="text-align: left">Consider that source for completion</td>
    </tr>
    <tr>
      <td style="text-align: left">No failure was recorded for that source</td>
      <td style="text-align: left">Keep it eligible for marking</td>
    </tr>
    <tr>
      <td style="text-align: left">The source passed this cache’s filter</td>
      <td style="text-align: left">Mark it in this cache</td>
    </tr>
  </tbody>
</table>

<p>The distinction between a source ID and a <code class="language-plaintext highlighter-rouge">(cache, source)</code> pair was the part I had to be most careful with. Deduplicating sources solves repeated writes after chunking. It doesn’t establish which branch did the work.</p>

<h2 id="retrying-can-repeat-successful-writes">Retrying can repeat successful writes</h2>

<p>An unhandled stream error exits <code class="language-plaintext highlighter-rouge">run()</code> before cache marking. A source that completed earlier in the same run can therefore be processed again next time. Storage and cache updates aren’t one transaction, so downstream writes still need appropriate retry behavior, such as idempotent persistence where the application requires it.</p>

<p>I’m willing to accept that repeated work here. It has a path to recovery. A premature cache entry can keep the unfinished source out of every later run unless someone invalidates it. For an indexing job, that is a much harder failure to notice from the query side.</p>

<p>The final cache writes use bounded concurrency rather than a serial round trip for every source. The trait also gains a required <code class="language-plaintext highlighter-rouge">NodeCache.set_by_id</code> method, so custom cache implementations need to support the deferred write. This is an API cost worth making visible when adopting the change.</p>

<p>The merged regressions cover ordinary success and failure, parent deduplication after chunking, partial chunk failure, and branch-specific completion. What I want to see in an indexing review now is the source’s retry path: if any stage loses part of this document, what state makes the next run pick it up again?</p>

<p>My <a href="https://github.com/bosun-ai/swiftide/pull/1164">Swiftide PR #1164</a> merged on September 16, 2026. <a href="https://github.com/bosun-ai/swiftide/pull/1164/files">Changes and regression tests</a> · <a href="https://github.com/bosun-ai/swiftide/blob/ad42fa97caee7e0702711222d28ec9c6f56c84d5/swiftide-indexing/src/pipeline.rs">Pipeline at the merge commit</a>.</p>]]></content><author><name>goingforstudying-ctrl</name></author><category term="ai-engineering" /><category term="retrieval" /><category term="rust" /><summary type="html"><![CDATA[An embedding failure should leave a document available for retry. Swiftide's cache had already decided it was done.]]></summary></entry><entry><title type="html">The test was protecting the bug</title><link href="https://goingforstudying-ctrl.github.io/notes/tensorflow-zero-gradient/" rel="alternate" type="text/html" title="The test was protecting the bug" /><published>2026-10-11T00:00:00+00:00</published><updated>2026-10-11T00:00:00+00:00</updated><id>https://goingforstudying-ctrl.github.io/notes/tensorflow-zero-gradient</id><content type="html" xml:base="https://goingforstudying-ctrl.github.io/notes/tensorflow-zero-gradient/"><![CDATA[<p>At <code class="language-plaintext highlighter-rouge">x = 0</code> and <code class="language-plaintext highlighter-rouge">y = 3.1</code>, TensorFlow’s <code class="language-plaintext highlighter-rouge">xlogy</code> returned a value of zero and a gradient of zero. The existing test agreed with both. The value was right. The gradient should have been about <code class="language-plaintext highlighter-rouge">1.1314</code>.</p>

<p>A passing test makes me hesitate in a way a failing one doesn’t. Before changing the implementation, I wanted an answer small enough to check without TensorFlow. Fix <code class="language-plaintext highlighter-rouge">y</code>, draw the function, and ask what happens just to either side of zero. The line goes straight through the origin. Its slope never disappears.</p>

<p>That left two things to change: the backward pass and the test that had been defending it.</p>

<h2 id="a-zero-value-can-still-carry-a-gradient">A zero value can still carry a gradient</h2>

<p>For positive <code class="language-plaintext highlighter-rouge">y</code>, <code class="language-plaintext highlighter-rouge">xlogy(x, y)</code> is <code class="language-plaintext highlighter-rouge">x * log(y)</code>. Holding <code class="language-plaintext highlighter-rouge">y = 3.1</code> makes it a straight line with slope <code class="language-plaintext highlighter-rouge">log(3.1)</code>. Setting <code class="language-plaintext highlighter-rouge">x</code> to zero picks a point on that line; it doesn’t change its slope.</p>

<div class="article-visual article-visual-gradient">
<figure class="gradient-figure" data-gradient="">
  <div class="figure-heading"><span class="eyebrow">Fig. 01 / TensorFlow</span><span class="figure-status"><span aria-hidden="true"></span>At the origin</span></div>
  <svg class="gradient-plot" viewBox="0 0 460 290" role="img" aria-labelledby="gradient-title gradient-description">
    <title id="gradient-title">A zero value with a nonzero slope</title>
    <desc id="gradient-description">The straight line f(x) = x times log(3.1) crosses the origin. Its slope is about 1.1314, including at zero. The old gradient incorrectly returned zero at that point.</desc>
    <defs><pattern id="plot-grid" width="32" height="32" patternUnits="userSpaceOnUse"><circle cx="1" cy="1" r="1" fill="currentColor" opacity=".16" /></pattern></defs>
    <rect x="28" y="12" width="404" height="258" rx="4" fill="url(#plot-grid)" />
    <g class="plot-axes" fill="none" stroke="currentColor" stroke-width="1"><path d="M40 154H422M214 24V270" /><path d="m416 150 6 4-6 4M210 30l4-6 4 6" /></g>
    <text class="plot-label" x="420" y="176">x</text><text class="plot-label" x="226" y="32">f(x)</text><text class="plot-label" x="195" y="176">0</text>
    <path class="correct-line" d="M58 254 370 54" fill="none" stroke-width="3" stroke-linecap="round" />
    <path class="old-tangent" d="M142 154H286" fill="none" stroke-width="2.5" stroke-dasharray="5 6" stroke-linecap="round" />
    <circle class="plot-halo" cx="214" cy="154" r="13" />
    <circle class="plot-point" cx="214" cy="154" r="5.5" stroke-width="3" />
    <text class="line-annotation" x="288" y="60">slope = log(3.1)</text><text class="old-annotation" x="246" y="188">old gradient: 0</text>
    <text class="equation-label" x="45" y="42">f(x) = x · log(3.1)</text>
  </svg>
  <div class="gradient-values" aria-live="polite" aria-atomic="true"><div><span>Forward value</span><strong data-forward="">0.0000</strong></div><div><span>Correct gradient</span><strong class="correct-value">1.1314</strong></div><div><span>Old gradient</span><strong class="old-value" data-old="">0.0000</strong></div></div>
  <div class="gradient-control" hidden=""><label for="gradient-x">Move x <output id="gradient-output" for="gradient-x">0.0</output></label><input id="gradient-x" type="range" min="-1" max="1" value="0" step="0.1" aria-describedby="gradient-hint" /><button type="button" class="text-button" data-reset="">Reset to zero ↺</button></div>
  <figcaption id="gradient-hint">The value is zero. The slope is still there.</figcaption>
</figure>

</div>

<p>Move the point away from zero and the old gradient agrees with the correct one. Reset it, and the error appears at the origin. This particular bug is easy to miss if a gradient check samples ordinary nonzero inputs.</p>

<p>The numerical check I kept coming back to needs only two nearby evaluations:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kn">from</span> <span class="nn">math</span> <span class="kn">import</span> <span class="n">isclose</span><span class="p">,</span> <span class="n">log</span>

<span class="n">y</span> <span class="o">=</span> <span class="mf">3.1</span>
<span class="n">h</span> <span class="o">=</span> <span class="mf">1e-6</span>
<span class="n">slope</span> <span class="o">=</span> <span class="p">((</span><span class="n">h</span> <span class="o">*</span> <span class="n">log</span><span class="p">(</span><span class="n">y</span><span class="p">))</span> <span class="o">-</span> <span class="p">(</span><span class="o">-</span><span class="n">h</span> <span class="o">*</span> <span class="n">log</span><span class="p">(</span><span class="n">y</span><span class="p">)))</span> <span class="o">/</span> <span class="p">(</span><span class="mi">2</span> <span class="o">*</span> <span class="n">h</span><span class="p">)</span>
<span class="k">assert</span> <span class="n">isclose</span><span class="p">(</span><span class="n">slope</span><span class="p">,</span> <span class="n">log</span><span class="p">(</span><span class="n">y</span><span class="p">),</span> <span class="n">rel_tol</span><span class="o">=</span><span class="mf">1e-12</span><span class="p">)</span>
</code></pre></div></div>

<p>This matters in a training graph because the backward pass sends a sensitivity upstream. If the loss contributes an upstream gradient <code class="language-plaintext highlighter-rouge">g</code>, this operation should return <code class="language-plaintext highlighter-rouge">g * log(y)</code> with respect to <code class="language-plaintext highlighter-rouge">x</code>, before any broadcast reduction. Returning zero erases that contribution. A forward-value check can pass while the optimizer receives the wrong signal.</p>

<h2 id="the-forward-shortcut-leaked-into-autodiff">The forward shortcut leaked into autodiff</h2>

<p>The old implementation was compact. It built a mask that was one wherever <code class="language-plaintext highlighter-rouge">x</code> was nonzero, then used <code class="language-plaintext highlighter-rouge">xlogy(mask, y)</code> as the partial derivative with respect to <code class="language-plaintext highlighter-rouge">x</code>.</p>

<p>I can see why that looked reasonable. At a nonzero <code class="language-plaintext highlighter-rouge">x</code>, the mask becomes one, and <code class="language-plaintext highlighter-rouge">xlogy(1, y)</code> gives <code class="language-plaintext highlighter-rouge">log(y)</code>. It also reuses the forward operation’s handling of zero. Unfortunately, that is exactly where the derivative gets lost: <code class="language-plaintext highlighter-rouge">xlogy(0, y)</code> deliberately returns zero.</p>

<p>The forward operation has a special value convention. Reusing it inside the gradient accidentally turned that convention into a statement about local sensitivity. Those are separate decisions, and the derivative only becomes obvious once you write them down separately.</p>

<p><code class="language-plaintext highlighter-rouge">xlog1py</code> used the same construction. Inside its logarithm domain, <code class="language-plaintext highlighter-rouge">y &gt; -1</code>, the function is <code class="language-plaintext highlighter-rouge">x * log1p(y)</code>, so the slope with respect to <code class="language-plaintext highlighter-rouge">x</code> is <code class="language-plaintext highlighter-rouge">log1p(y)</code> at the origin too.</p>

<p>The patch uses the logarithms directly:</p>

<table>
  <thead>
    <tr>
      <th style="text-align: left">Operation</th>
      <th style="text-align: left">Partial with respect to <code class="language-plaintext highlighter-rouge">x</code></th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td style="text-align: left"><code class="language-plaintext highlighter-rouge">xlogy(x, y)</code></td>
      <td style="text-align: left"><code class="language-plaintext highlighter-rouge">log(y)</code></td>
    </tr>
    <tr>
      <td style="text-align: left"><code class="language-plaintext highlighter-rouge">xlog1py(x, y)</code></td>
      <td style="text-align: left"><code class="language-plaintext highlighter-rouge">log1p(y)</code></td>
    </tr>
  </tbody>
</table>

<p>The surrounding gradient machinery still multiplies by the upstream gradient, reduces broadcast dimensions, and reshapes the result to the input shape. I left the <code class="language-plaintext highlighter-rouge">y</code> partials alone. A correction to one derivative shouldn’t acquire unrelated changes to broadcasting or the other input’s gradient.</p>

<h2 id="the-regression-needed-an-independent-answer">The regression needed an independent answer</h2>

<p>Changing the code made the existing zero-<code class="language-plaintext highlighter-rouge">x</code> expectations wrong. I changed those expectations to the logarithms and kept the checks that the <code class="language-plaintext highlighter-rouge">y</code> gradients are zero for these inputs. The regression now distinguishes the two partials instead of letting a shared zero hide the difference.</p>

<p>The merged tests exercise <code class="language-plaintext highlighter-rouge">float16</code>, <code class="language-plaintext highlighter-rouge">float32</code>, and <code class="language-plaintext highlighter-rouge">float64</code>. They also cover <code class="language-plaintext highlighter-rouge">y = 0</code> for <code class="language-plaintext highlighter-rouge">xlogy</code> and <code class="language-plaintext highlighter-rouge">y = -1</code> for <code class="language-plaintext highlighter-rouge">xlog1py</code>, where TensorFlow’s implemented <code class="language-plaintext highlighter-rouge">x</code> gradient is negative infinity. These boundary cases have no finite classical derivative. The straight-line argument above applies inside the logarithm domain; the boundary tests pin down the framework’s chosen behavior.</p>

<p>I care about that distinction because a numerical check and an API boundary test answer different questions. At <code class="language-plaintext highlighter-rouge">y = 3.1</code>, I can derive a finite slope and compare it with nearby values. At the singularity, I need the test to state the intended behavior explicitly.</p>

<p>This patch changed only a few expressions, but it changed what a green test meant to me. When a gradient expectation looks suspicious, I now want a derivation or an independent numerical check beside it. Reproducing the implementation in the expected value can preserve the same mistake for years.</p>

<p>My <a href="https://github.com/tensorflow/tensorflow/pull/119869">TensorFlow PR #119869</a> merged on July 16, 2026. <a href="https://github.com/tensorflow/tensorflow/pull/119869/files">Implementation and test changes</a>.</p>]]></content><author><name>goingforstudying-ctrl</name></author><category term="machine-learning" /><category term="numerical-correctness" /><summary type="html"><![CDATA[TensorFlow returned a zero gradient, and an existing test agreed. A straight line through the origin gave me a reason to doubt them both.]]></summary></entry><entry><title type="html">Why this blog exists</title><link href="https://goingforstudying-ctrl.github.io/2026/08/27/why-this-blog.html" rel="alternate" type="text/html" title="Why this blog exists" /><published>2026-08-27T00:00:00+00:00</published><updated>2026-08-27T00:00:00+00:00</updated><id>https://goingforstudying-ctrl.github.io/2026/08/27/why-this-blog</id><content type="html" xml:base="https://goingforstudying-ctrl.github.io/2026/08/27/why-this-blog.html"><![CDATA[<p>I use this blog to explain changes I contribute to open source. The topics include numerical correctness in ML frameworks, indexing and inference systems, and concurrency in distributed backends.</p>

<p>The diff shows what changed. I want the write-up to explain why that change is enough, which test would have caught the failure, and what tradeoffs remain.</p>

<p>Each technical post links to its merged PR so readers can check the implementation and tests.</p>]]></content><author><name>goingforstudying-ctrl</name></author><summary type="html"><![CDATA[I use this blog to explain changes I contribute to open source. The topics include numerical correctness in ML frameworks, indexing and inference systems, and concurrency in distributed backends.]]></summary></entry></feed>