Insider · Optimus · bugpredict5_5

Bir QA task'ı, gerçekten bug mu? A QA ticket — is it really a bug?

bugpredict5_5, Jira'ya düşen her QA task'ını okuyup dört soruyu cevaplıyor: bug mu, kök neden ne, kim çözer ve mümkünse çözümün kendisi. Aşağıdaki konsol, bunu adım adım — gerçek task'lar ve gerçek çıktılarla — gösteriyor. bugpredict5_5 reads every QA ticket that lands in Jira and answers four questions: is it a bug, what is the root cause, who fixes it, and where possible the actual solution. The console below walks through it step by step — with real tickets and real outputs.

Bu sayfadaki her örnek, bugpredict5_5/ altındaki gerçek günlük raporlardan alındı (6–30 Temmuz 2026, 17 koşu günü, 141 task). Hiçbir çıktı uydurulmadı. Every example on this page is taken from the real daily reports under bugpredict5_5/ (6–30 July 2026, 17 run days, 141 tickets). No output here is invented.


Ne üretiyor What it produces

Her task için dört cevapFour answers per ticket

Çıktı bir puan ya da bir etiket değil. Geliştiricinin, QA'in veya partnerin olduğu gibi kullanabileceği dört cümle. The output is not a score or a tag. It is four sentences a developer, a QA or a partner can use as-is.

1

Bug mu?Is it a bug?

Beş kategoriden biri — ve o kategoriye hangi kanıt sınıfının taşıdığı. One of five categories — plus which evidence class carried it there.

2

Kök neden ne?What is the root cause?

Mekanizma, düz dille. "Şuraya bak" değil, "şu yüzden oluyor". The mechanism, in plain language. Not "look here", but "this is why it happens".

3

Kim çözer?Who fixes it?

Ürün ekibi mi, Optimus mu, partner mı — yoksa hiç kimse mi (beklenen davranış)? A product team, Optimus, the partner — or nobody, because it is expected behaviour?

4

ÇözümThe solution

Araştırma cevaba ulaştıysa: kopyalanıp yapıştırılabilir çözüm metni. When the research actually reaches an answer: copy-pasteable resolution text.


Pipeline Pipeline

Bir task adım adım nasıl işleniyorHow one ticket is processed, step by step

Önce bir task seç, sonra Oynat'a bas. Sol tarafta o adımda ne yapıldığı düz dille, sağ tarafta o adımın gerçek çıktısı görünür. Kesikli halkalar atlanan adımlardır — her task her adımı çalıştırmaz, bu kasıtlı. Pick a ticket, then hit Play. The left panel says what happens at that step in plain language; the right panel shows that step's real output. Dashed rings are skipped steps — not every ticket runs every step, and that is by design.

bugpredict5_5 · örnek koşusample run
READY

    telemetry— idle
    not:note: Buradaki adımlar ve çıktılar gerçek günlük raporlardan alınmıştır; canlı bir koşu akışı değil, birebir yeniden oynatımdır. These steps and outputs are taken from the real daily reports — this is a faithful replay, not a live job tail.

    KategorilerCategories

    Beş cevap, beş farklı anlamFive answers, five different meanings

    Sayılar 6–30 Temmuz 2026 arasındaki 17 koşu gününden ve 141 tasktan geliyor. Her task'ın en son kararı sayıldı. The counts come from 17 run days between 6–30 July 2026 and 141 tickets. Each ticket's most recent verdict is counted.

    1TASKTICKET
    Cat 1

    Kesin product bug. Kodun kendisi okundu, hangi satırın semptomu ürettiği gösterildi ve düşman avukatı bunu çürütemedi. Ürün ekibine escalate edilir. Confirmed product bug. The code itself was read, the exact lines producing the symptom were shown, and the hostile reviewer could not refute it. Escalated to the product team.

    Örnek: OPT-236634 — editörün "geçerli alanlar" listesi ağacı tek seviye geziyor. Bu eşiğe 141 task içinde yalnızca 2 task ulaştı; biri ertesi gün Cat 1.5'e kalibre edildi (bulgu korundu, güven düşürüldü). Example: OPT-236634 — the editor's "valid fields" list walks the tree one level deep. Only 2 tickets out of 141 ever reached this bar; one was recalibrated to Cat 1.5 the next day (finding preserved, confidence lowered).
    14TASKTICKETS
    Cat 1.5

    Güçlü şüphe — kesin DEĞİL. Kanıt gerçek ama kod nedenselliği okunmamış ya da tek bir alternatif elenmemiş. Her zaman "neden kesin değil" ve "hangi 1–3 adım kesinleştirir" ile birlikte gelir. Strong suspicion — NOT confirmed. The evidence is real but the code causation was not read, or one alternative was not excluded. Always ships with "why not confirmed" and "which 1–3 steps would settle it".

    Buraya iki ayrı yoldan gelinir. Kanıt eksikse: OPT-238009 — panelde 5 koşullu segment, gönderimde 3 koşul; kod okunmadığı için 0.82. Kanıt varken bile: OPT-237276 — kod okundu, nedensellik kanıtlandı, canlı test ayrımı doğruladı; ama düşman avukatının tek bir alternatifi ("partner o entegrasyonu hiç kurmamış olabilir") elenemediği için 0.85 ile yine Cat 1.5. Kod kanıtı tek başına "kesin" demeye yetmiyor. There are two separate routes here. When evidence is missing: OPT-238009 — a segment stored with 5 conditions ran with 3; the code was never read, so 0.82. Even when evidence exists: OPT-237276 — the code was read, causation was proven, a live test confirmed the separation; but one alternative from the hostile reviewer ("the partner may never have set up that integration") could not be eliminated, so it is still Cat 1.5 at 0.85. Code evidence alone is not enough to say "confirmed".
    55TASKTICKETS
    Cat 2

    Doğrulama gerekiyor. Kapatmaya da bug demeye de yetecek kanıt yok. Kritik kısım şu: bu kategoride kalan her task, onu çevirecek tek kanıtın adıyla birlikte yazılır — "araştırılmalı" demekle yetinilmez. Needs verification. There is not enough evidence to close it or to call it a bug. The critical part: every ticket left here is written up with the name of the single piece of evidence that would flip it — never a vague "needs investigation".

    Örnek: OPT-238100 — "Cat 1.5'e çevirecek tek kanıt: push alan somut kullanıcılarda push_delivered olayının var olduğunu gösterip panelde Delivered'ın 0 kaldığını yan yana koymak." Example: OPT-238100 — "The one piece of evidence that flips this: show that push_delivered exists on concrete users who received the push, next to a panel still reading Delivered 0."
    70TASKTICKETS
    Cat 3

    Bug değil. Partner konfigürasyonu, veri, beklenen davranış ya da Optimus'un kendi işi. Ama "bug değil" demek yeterli sayılmıyor: bu kategorinin asıl çıktısı cevabın kendisi. Not a bug. Partner configuration, data, expected behaviour, or Optimus' own work. But "not a bug" is not accepted as an answer on its own: the real output of this category is the answer itself.

    Örnek: OPT-238018 — "bozuk UUID" sanılan kimlik, dokümante anonim kullanıcı biçimi; zaman damgası bağımsız olarak kampanya giriş tarihini doğruluyor. Partnera gidecek cevap hazır yazıldı. Example: OPT-238018 — the "corrupted UUID" is the documented anonymous-user format; its timestamp independently confirms the campaign entry date. The reply for the partner was written out.
    1TASKTICKET
    Cat 4

    Yetersiz açıklama. Beş temel öğeden ikisi birden eksikse araştırma başlamaz — tek bir model çağrısı harcanır ve raporlayana yapıştırılabilir bir eksik listesi çıkarılır. Not enough information. If two of the five basic elements are missing, research does not start — one model call is spent and a paste-ready list of what is missing goes back to the reporter.

    Bu kategori nadir çünkü yorumlarda sonradan verilen bilgiler de sayılıyor; sırf açıklama kısa diye task geri çevrilmiyor. This category is rare because facts supplied later in comments count too; a ticket is never bounced merely for having a short description.
    Cat 1 · 1 Cat 1.5 · 14 Cat 2 · 55 Cat 3 · 70 Cat 4 · 1

    BekçilerThe guards

    Neden sürekli "bug!" diye bağırmıyorWhy it does not keep crying wolf

    Bir task'a "bug" diyebilmek için sekiz kanıt sınıfından en az birinin ayakta kalması gerekiyor — ve açıklamadaki çelişkiye dayanan sınıf, 23 ayrı bekçi kuralından geçmek zorunda. Aşağıdaki gerçek örnek, bu bekçilerin ne işe yaradığını en net gösteren vaka. To call a ticket a bug, at least one of eight evidence classes must survive — and the class based on a contradiction in the description must pass 23 separate guard rules. The real case below is the clearest demonstration of what those guards are for.

    az kalsın yanlış alarmnearly a false alarm

    OPT-238088 · trendyolyemekb2b

    2026-07-30

    Şikâyet kusursuz bir bug gibi duruyordu: "Profilde var olan attribute, kampanya değerlendirmesinde bulunamıyor." Bu, iki Insider ekranının birbirini tutmadığı klasik durum — tam olarak "bug" demeye yeten imza. The complaint looked like a textbook bug: "An attribute that exists on the profile is not found during campaign evaluation." That is the classic case of two Insider surfaces disagreeing — exactly the signature that justifies calling something a bug.

    Bekçi kuralı devreye girdi: karşılaştırmanın iki tarafı aynı anahtarı taşıyor mu? Panelin bulamadığı ad ..._claims_cost, profildeki adlar ise ..._claims_count. Farklı adlar. Yani ortada çelişki yok — sadece kampanyada var olmayan bir alan adı yazılmış. A guard rule fired: do both sides of the comparison carry the same key? The name the panel could not find is ..._claims_cost, while the profile holds ..._claims_count. Different names. So there is no contradiction — the campaign simply references a field that does not exist.

    Raporun kendi cümlesi: "İlk 1–2 görselle durulsa bu ticket Cat 1.5 olarak işaretlenecekti." 9 ekten 7'si kronolojik sırayla okunduğu ve canlı profil kontrol edildiği için yanlış alarm engellendi — ve yerine partnerın bugün uygulayabileceği gerçek çözüm çıktı. The report's own words: "Had this stopped at the first 1–2 images, the ticket would have been flagged Cat 1.5." Because 7 of 9 attachments were read in chronological order and the live profile was checked, the false alarm was prevented — and replaced by a fix the partner can apply today.
    kural
    rule

    Yakınlık ≠ nedensellikAdjacency ≠ causation

    Konuyla ilgili görünen kod yetmez. O kodun, tarif edilen girdide semptomu gerçekten ürettiği gösterilmek zorunda. Gösterilemezse bulgu bir kademe düşer. Code that merely looks related is not enough. It must be shown to actually produce the symptom on the described inputs. If not, the finding drops a rung.

    kural
    rule

    Zaman örtüşmesi kanıt değildirTime overlap is not evidence

    "Deploy oldu, sonra sorun başladı" tek başına hiçbir zaman yük taşıyamaz. Mekanizma bağı yoksa sadece destekleyici nottur. "A deploy happened, then the problem started" can never be load-bearing on its own. Without a mechanism link it is only a corroborating note.

    kural
    rule

    Doğru çalışan yedek plan bug değildirA working fallback is not a bug

    Kod bir "yedek/varsayılan" dala düşüyorsa asıl soru neden o dala düştüğü. Tetikleyen şey partnerın gönderdiği veriyse, kod doğru çalışıyordur. If the code lands on a fallback/default branch, the real question is why it was triggered. If the trigger is partner input, the code is working correctly.

    kural
    rule

    Bileşik task'ta ana şikâyet bağlanmalıIn a compound ticket, bind the core

    Bir task'ta üç şikâyet varsa, yan şikâyetlerden birine bulunan kanıt bütün task'ı "bug" yapmaz. Kanıt ana şikâyete bağlanmak zorunda. If a ticket carries three complaints, evidence for a side complaint does not make the whole ticket a bug. The evidence must bind the core symptom.


    Ölçülen sonuçMeasured results

    "Bug" dediğinde neden gerçekten bug çıkıyor Why it is usually right when it says "bug"

    24 Haziran 2026 denetiminde, üretimde işaretlenmiş kayıtlar gerçek sonuçlarıyla karşılaştırıldı — rubrikle değil, task'ın gerçekten nasıl bittiğiyle. In the 24 June 2026 audit, the records flagged in production were compared against their actual outcomes — not against a rubric, but against how each ticket really ended.

    %95,2
    30 günlük pencerede doğru tahmincorrect predictions, 30-day window
    21 kayıttanrecords → 20
    %90,0
    50 günlük pencerede doğru tahmincorrect predictions, 50-day window
    40 kayıttanrecords → 36
    13/17
    ürün panosuna taşınan task, gerçek ürün hatası olarak çözüldü/yayınlandı tickets moved to a product board were resolved/released as real product bugs
    +2 aktif incelemede · yalnız 2 by-design+2 in active review · only 2 by-design
    689 → 43
    taranan task → işaretlenen kayıttickets scanned → records flagged
    %6'lık son derece seçici bir işaretlemea highly selective 6% flag rate
    Bu rakam kime ait, dürüstçe: denetim, üretimdeki potential-bug-ai etiketini Mayıs–Haziran 2026 penceresinde ölçtü — yani bugpredict hattının o dönemki sürümünü. bugpredict5_5 (6 Temmuz'da başladı) bu makineyi devraldı: kendi dosyasında her kapının hangi sürümden geldiğini gösteren bir köken haritası ve denetimin kaçırdıklarını içeren bir regresyon seti taşıyor. 5_5'in kendi precision'ı geriye dönük olarak doğrulandı, ileriye dönük olarak henüz ölçülmedi. Örneklem küçük — 40 kayıtta güven aralığı yaklaşık %77–96. Whose number this is, honestly: the audit measured the production potential-bug-ai label over a May–June 2026 window — i.e. the version of the bugpredict line running then. bugpredict5_5 (which started on 6 July) inherited that machinery: its own file carries a provenance map naming which version each gate came from, plus a regression set built from what the audit missed. 5_5's own precision is validated retrospectively and has not yet been measured prospectively. The sample is small — at 40 records the confidence interval is roughly 77–96%.
    en güçlü vakathe strongest case

    OPT-234410 → SD-144300 · SIA · Byte Masters

    A/B testinde gösterilen ortalama sepet tutarı anormaldi. İki ayrı yapay zekâ sistemi ve ürün ekibinin QA'i — üçü de "bug yok" dedi. bugpredict kodu birebir okudu: hesaplanan bir değer hiç döndürülmüyordu. Asıl kanıt ise raporun kendi içindeki matematiksel çelişkisiydi — ekranda gösterilen %92,16'lık artış yalnızca bir sayıdan türeyebilirdi, ama yan hücrede bambaşka bir sayı vardı. Byte Masters kararını değiştirdi ve Product Bug açıldı. The average order value shown in an A/B test was anomalous. Two separate AI systems and the product team's own QA — all three said "no bug". bugpredict read the code line by line: a value was being computed and then never returned. The decisive evidence was the report's own internal mathematical contradiction — the 92.16% uplift on screen could only derive from one number, yet the adjacent cell showed a completely different one. Byte Masters reversed its decision and a Product Bug was opened.

    Bilgi-tabanı AIKnowledge-base AI Slack AISlack AI bugpredict
    SonuçConclusion "Tasarım gereği — fark beklenen." → bug değil"By design — the difference is expected." → not a bug "Uyuşmazlık farklı formülden gelmiyor." → bug değil yönünde"The mismatch is not from a different formula." → leaning not-a-bug "Değer hesaplanıp hiç döndürülmüyor; rapor kendiyle çelişiyor." → gerçek bug"The value is computed and never returned; the report contradicts itself." → a real bug
    DayanağıBasis Genel metodoloji açıklamasıA general methodology explanation Kod özeti — ama iç çelişkiyi kaçırdıA code summary — but it missed the internal contradiction dosya:satır kod okuması + matematiksel iç çelişki + canlı sayı doğrulaması + aynı dosyadaki emsal düzeltmefile:line code read + the mathematical contradiction + live number verification + a precedent fix in the same file

    Aradaki fark neydi: diğer iki sistem makul ve akıcı açıklamalar üretti — ama ikisi de "ürünün kendi çıktısı kendiyle çelişiyor" tespitini yapamadı, çünkü tek atışta özet üretiyorlar; kodu canlı sayılarla ve raporun iç tutarlılığıyla çapraz doğrulamıyorlar. bugpredict "ilgili görünen" açıklamayı değil matematiksel olarak imkânsız olanı arıyor. Bunlar rakip değil tamamlayıcı araçlar — ama "escalate edelim mi" kararında belirleyici olan disiplinli hat. What made the difference: the other two systems produced plausible, fluent explanations — but neither could make the observation that "the product's own output contradicts itself", because they generate one-shot summaries; they do not cross-verify code against live numbers and the report's internal consistency. bugpredict does not look for the explanation that sounds related — it looks for what is mathematically impossible. These are complementary tools, not competitors — but on the "should we escalate" decision, the disciplined path is what decides.

    Bu sonucu mümkün kılan yeteneklerThe capabilities behind that result

    Hiçbiri tek başına yeterli değil — güç, dördünün sırayla uygulanmasından geliyor. None of these is sufficient alone — the strength comes from applying all four in order.

    1
    Önce bağlamı kurarFirst it establishes context
    kim sahibi · ne dökümante · görsellerde ne var · geçmişte ne olduwho owns it · what is documented · what the images show · what happened before
    Sahipliği en başta çözer.It resolves ownership first. Ticket'taki ürün/partner/belirtiyi hangi ekibin hangi reposuna ait olduğuna bağlar. Sonraki tüm muhakeme bu sahiplik üzerine kurulur — "kime escalate" kararının temeli budur. Eşleme güvenli değilse kayıt temkinli işlenir. It maps the product/partner/symptom to a specific team's repository. All later reasoning rests on that ownership — it is the foundation of the "who do we escalate to" decision. When the mapping is not confident, the record is handled conservatively.
    Kendi kodunu asla başka ekibe atmaz.It never pushes its own code onto another team. Hata Optimus'un kendi kodundaysa bunu bir ürün ekibine yönlendirmez — çünkü "ürün hatası" demek tanımı gereği "başka bir ekibe taşı" demektir. Ürün ekiplerine yanlış yönlendirme gitmesini engeller; güven için kritik. If the defect is in Optimus' own code it is never routed to a product team — because "product bug" means, by definition, "move this to another team". This prevents misrouted escalations reaching product teams; critical for trust.
    Görselleri "var" diye saymaz, gerçekten açar.It does not count images as present — it actually opens them. Jira eklerini ve dış bağlantıları indirip görüntüler: hangi ekran, görünen değerler, hata banner'ı, konsol çıktısı. Bu kanıt sonradan kod doğrulama ve düşman denetçi aşamalarına orijinal pikseliyle geri beslenir. Böylece "ekran görüntüsü gönderin" tuzağına düşmez — kanıt zaten ekteyse onu kullanır. It downloads and views Jira attachments and external links: which screen, the rendered values, the error banner, the console output. That evidence is later fed back into code verification and the hostile reviewer as the original pixels. So it never falls into the "please send a screenshot" trap — if the evidence is already attached, it uses it.
    Geçmiş vakalardan öğrenir.It learns from past cases. Hem eski çözümleri hem eski ret gerekçelerini kanıt olarak çeker — biri kaçırmayı azaltır, diğeri yanlış alarmı. Tekerleği yeniden icat etmez. It pulls both past resolutions and past rejection reasons as evidence — the first reduces misses, the second reduces false alarms. It does not reinvent the wheel.
    2
    Sonra kanıt toplarThen it gathers evidence
    niyet · kod · son değişiklikler · gözlemlenebilirlikintent · code · recent changes · observability
    Soruyu bug'dan ayırır.It separates a question from a bug. Ucuz bir ön eleme: bu gerçek bir hata raporu mu, yoksa bir soru / araştırma talebi mi? Soru niteliğindeki kayıtlar otomatik olarak "doğrulama gerekiyor" tavanına çekilir. "Şunu açıklar mısınız" bug sayılmaz. A cheap pre-filter: is this a genuine defect report, or a question / investigation request? Question-shaped records are automatically held at the "needs verification" ceiling. "Can you explain this" is never counted as a bug.
    Kök nedeni dosya:satır düzeyinde gösterir.It shows the root cause at file:line level. Sadece "muhtemelen bug" demez; ilgili repoyu, dosyayı ve satırı verir. Ürün ekibi karşısına tartışmaya hazır kanıtla çıkar — bu, escalate edilen kayıtların kabul görme oranını doğrudan yükselten şey. It never stops at "probably a bug"; it names the repository, the file and the line. It arrives in front of the product team with evidence ready to be argued — which is exactly what raises the acceptance rate of what it escalates.
    Zaten yayınlanmış düzeltmeyi tekrar escalate etmez.It does not re-escalate a fix that already shipped. Açık bir ürün kaydı varsa kaçırmaz; ama düzeltme çoktan yayınlanmışsa yeniden açmaz — partnera kök nedeni iletir. Ürün ekiplerinin zamanını koruyan sessiz ama önemli bir davranış. If an open product record exists it will not be missed; but if the fix already shipped it is not reopened — the root cause goes back to the partner instead. A quiet but important behaviour that protects product teams' time.
    Gözlemlenebilirlik sinyalini kod kanıtıyla birleştirir.It combines observability signals with code evidence. Belirsiz bir "bazen zaman aşımı alıyoruz" şikâyeti, yük pencerelerindeki gecikme sıçramalarıyla eşleştirilince ürün tarafı altyapı sorununa indirgenebiliyor — ve partner konfigürasyonu kanıtla eleniyor. A vague "we sometimes get timeouts" complaint, matched against latency spikes inside specific load windows, reduces to a product-side infrastructure problem — and partner configuration gets ruled out with evidence.
    3
    Sonra kodu doğrularThen it verifies the code
    parçala · nedensellik · tasarım niyeti · çelişkidecompose · causation · design intent · contradiction
    Şikâyeti parçalara böler, her parçayı ayrı doğrular.It breaks the complaint into parts and verifies each separately. Bir task'ı "bütün olarak bug / değil" diye yargılamaz. Her parça bağımsız bir doğrulama zincirinden geçer ve ayrı sınıflandırılır. Böylece "kod bozuk görünüyor" toptan yargısı yerine hangi parçanın gerçekten kırık olduğu netleşir. It never judges a ticket as "a bug / not a bug" as a whole. Each part goes through an independent verification chain and is classified separately. Instead of a blanket "the code looks broken", which part is actually broken becomes clear.
    "İlgili" değil, "neden olan" testi.The test is "causes it", not "relates to it". Kodun konuyla ilgili görünmesi yetmez; o kodun bu hatayı gerçekten ürettiğini zorunlu kılar. Yüzeyde eşleşen ama aslında bozuk olmayan kod yollarını eler. Yanlış alarmın en büyük tek kaynağı budur ve doğrudan hedef alınır. It is not enough for the code to look related; it must be shown to actually produce the failure. Code paths that match on the surface but are not broken get eliminated. This is the single largest source of false alarms, and it is targeted directly.
    Partner verisine bağlı belirtileri ürün hatası saymaz.Symptoms caused by partner data are not counted as product bugs. Yalnızca eksik veya bozuk partner verisiyle ortaya çıkan belirtiler bir tavana takılır — ve bu tavan aynı zamanda partnera verilecek çözümü de üretir. Symptoms that only appear with missing or malformed partner data hit a ceiling — and that same ceiling also produces the answer to hand back to the partner.
    "Tasarım gereği" iddiası ancak çürütülemiyorsa kabul edilir.A "by design" claim is accepted only if nothing refutes it. Kod dökümante edilmiş kasıtlı davranışla eşleşiyorsa karar "bug değil" tabanına çekilir — ama herhangi bir kanıt bu davranışı ihlal ediyorsa kapı açılır. Yıldızlı vakada tam olarak bu işledi: iki sistem "tasarım gereği" derken, raporun kendi iç çelişkisi bu kapıyı açtı ve bug doğrulandı. If the code matches documented intentional behaviour, the verdict drops to the "not a bug" floor — but if any evidence violates that behaviour, the door opens. That is exactly what happened in the starred case: two systems said "by design", and the report's own internal contradiction opened the door and confirmed the bug.
    4
    En son karar verirOnly then does it decide
    tartışma · tutarlılık · sahip · güven eşiğidebate · consistency · owner · confidence floor
    Savcı ⇄ savunma ⇄ hakim.Prosecutor ⇄ defence ⇄ judge. Her aday için iki karşıt görüş birbiriyle tartışır: biri "bu bir ürün hatası" diye iddia eder, diğeri dökümante davranış / emsal task / partner konfigürasyonu ile çürütmeye çalışır. Hakem iki tarafı birebir alıntılayarak karar verir. Taraflar ayrışırsa daha temkinli kategori seçilir. For every candidate, two opposing positions argue against each other: one asserts "this is a product bug", the other tries to refute it with documented behaviour, a precedent ticket, or partner configuration. The judge decides by quoting both sides verbatim. When they diverge, the more conservative category wins.
    Düşman denetçi yalnızca düşürebilir, asla yükseltemez.The hostile reviewer can only demote, never promote. Tek yönlü bir kapı: bir iddia kendini savunamazsa elenir, ama hiçbir denetçi bir kaydı yukarı taşıyamaz. Yanlış pozitifleri ve onay yanlılığını asıl eleyen aşama budur — ve tam da bu asimetri, "bug" dendiğinde ona güvenilebilmesini sağlıyor. A one-way door: a claim that cannot defend itself is eliminated, but no reviewer can ever move a record up. This is the stage that actually removes false positives and confirmation bias — and it is precisely that asymmetry that makes "bug" worth trusting.
    "Güçlü Şüphe" ara katmanı kaçırmayı önler.The "Strong Suspicion" middle tier prevents misses. Kararı "kesin bug / bug değil" diye ikiye sıkıştırmaz. Kök neden tek satıra indirilemese bile güçlü işaret varsa kayıt kesin denmeden ama kaçırılmadan işaretlenir — zorunlu doğrulama checklist'iyle. İkili bir sınıflandırıcının ya yanlış "kesin" diyeceği ya da gözden kaçıracağı gerçek hatalar burada yakalanıyor. It does not compress the decision into "confirmed bug / not a bug". Even when the root cause cannot be pinned to a single line, a strong signal gets flagged without claiming certainty and without being dropped — with a mandatory verification checklist. Real bugs that a binary classifier would either wrongly call "confirmed" or quietly lose are caught here.
    Emin değilse emin gibi davranmaz.When it is not sure, it does not act sure. Her karar bir güven skoru taşır ve eşiği geçmeyen aday otomatik olarak bir kademe aşağı iner. Ayrıca son bir kapı: "bunu kim düzeltir?" — sahibin gerçekten başka bir ürün ekibi olduğu kanıtlanmadan "ürün hatası" denmiyor. Every decision carries a confidence score, and a candidate below the floor automatically drops a rung. Plus one final gate: "who fixes this?" — nothing is called a product bug until the owner is proven to genuinely be another product team.
    hataları teste dönüştürmekturning misses into tests

    Denetimin kaçırdıkları artık birer regresyon testi What the audit got wrong is now a regression test

    Denetim 4 yanlış tahmin buldu. Bunların üçü — bir eski kampanya kuralı kimliği, bir tekrarlanamayan geçici hata, ve kullanıcılarda telefon alanı olmadığı için oluşan bir API hatası — bugün bugpredict5_5'in dosyasında adı konmuş birer bekçiye bağlı regresyon vakası olarak duruyor. Yani her biri, "bu bir daha olursa hangi kural onu durduracak" sorusuna yazılı bir cevapla eşleşiyor. The audit found 4 wrong predictions. Three of them — a stale campaign rule id, an unreproducible transient error, and an API failure caused by users missing a phone field — now sit in bugpredict5_5's own file as regression cases bound to a named guard. Each one is matched to a written answer to the question "if this happens again, which rule stops it?"

    Bir sistemin kalitesini asıl gösteren şey, doğru bildikleri değil — yanlış bildiklerini nasıl kalıcı hâle getirdiğidir. Toplam 12 geçmiş yanlış alarm, bugün her değişiklikte korunması gereken bir kontrol listesi. What really shows a system's quality is not what it got right — it is what it does with what it got wrong. Twelve historical false alarms are now a checklist that every future change has to keep satisfying.

    HafızaMemory

    Yarın yeni bir kanıt gelirse ne oluyor?What happens when new evidence arrives tomorrow?

    Her karar hafızaya yazılıyor. Aynı task tekrar karşıya geldiğinde eski karar okunuyor — ve onu değiştirmek için tek bir kural var. Every decision is written to a ledger. When the same ticket comes back, the old decision is read — and there is exactly one rule for changing it.

    14 → 15 Temmuz

    OPT-236634

    Cat 1 → Cat 1.5 · bulgu korunduCat 1 → Cat 1.5 · finding preserved

    14 Temmuz'da task "kesin bug" olarak işaretlendi. Ertesi gün, aynı partnerda deneyimli bir QA taze bir journey'de hatayı tekrar edemedi. On 14 July the ticket was marked "confirmed bug". The next day an experienced QA on the same partner could not reproduce the error on a fresh journey.

    Kolay olan şey bulguyu geri çekmek olurdu. Kural buna izin vermedi: tekrar edememek mekanizmayı çürütmüyor, sadece onunla yarışıyor. Çürütme yükü karşılanmadığı için bulgu, kanıt sınıfı, sorumlu ekip ve escalate tavsiyesi aynen korundu. The easy move would have been to withdraw the finding. The rule did not allow it: a non-reproduction does not refute the mechanism, it merely competes with it. Because the burden of refutation was not met, the finding, its evidence class, the owning team and the escalate recommendation were all preserved.

    Değişen tek şey güven etiketi oldu: "kesin" yerine "güçlü şüphe + tek doğrulama adımı". Yeni gözlem silinmedi — rakip hipotez olarak rapora yazıldı. The only thing that changed was the confidence label: "confirmed" became "strong suspicion + one verification step". The new observation was not deleted — it was written into the report as a competing hypothesis.
    30 Temmuz koşusuJuly run

    Gecikmiş kanıt taramasıThe late-evidence scan

    8 gecikmiş sinyal bulundu8 late signals found

    En büyük kaçırma sebebi model değil, zamanlama: belirleyici yorum çoğu zaman değerlendirmeden sonra geliyor. Bu yüzden son 14 günün "bug değil" denmiş task'ları her gün yeniden taranıyor — sadece değişen kısım okunarak, hiç model çalıştırılmadan. The biggest source of misses is not the model but timing: the decisive comment usually lands after the evaluation. So the last 14 days of "not a bug" tickets are re-scanned every day — reading only the delta, with no model running at all.

    30 Temmuz koşusunda bu tarama, daha önce Cat 2 denmiş 8 taskın sonradan sert sinyal kazandığını buldu: ikisi "Blocked (Product)" durumuna geçmiş, altısı bir ürün bug kaydına bağlanmış. On the 30 July run this scan found 8 tickets previously called Cat 2 that had since gained a hard signal: two had moved to "Blocked (Product)", six had been linked to a product bug record.

    Hiçbiri geriye dönük etiketlenmedi. Hepsi öğrenme kaydı olarak yazıldı — çünkü geçmişi düzeltmek değil, gelecekteki kaçırmayı azaltmak hedefleniyor. None of them were retro-labelled. All were written down as learning records — the goal is not to rewrite the past but to reduce the next miss.

    SayılarNumbers

    17 koşu günü, 141 task17 run days, 141 tickets

    Bir task'a harcanan iş, cevabın zorluğuna göre değişiyor. Bu kasıtlı: kolay task'ta pahalı doğrulama makinesi hiç çalışmıyor. The work spent on a ticket scales with how hard the answer is. That is deliberate: on an easy ticket the expensive verification machinery never runs at all.

    86%
    somut cevapla kapandı
    (çözüm üretildi veya sahibi adıyla yönlendirildi)
    closed with a concrete answer
    (solution produced or owner named)
    43
    task'ta çözümün kendisi yazıldıtickets where the solution itself was written
    78
    task net sahibine yönlendirilditickets routed to a named owner
    20
    task açık kaldı (17 açık + 3 bilgi eksik)tickets left open (17 open + 3 need info)
    Cat 4
    1
    model çağrısı — araştırma başlamadan dururmodel call — stops before research
    hemen bitirfinish now
    3
    çağrı — kural zaten kararı veriyor, doğrulama harcanmıyorcalls — the rules already decide; no verification spent
    doğrulamaverification
    3–5
    çağrı — izleme listesi + tek kanıt tanımıcalls — watchlist + naming the flip evidence
    bug adayıflag candidate
    5–7
    çağrı — düşman avukatı + hakem burada devreye girercalls — this is where the hostile reviewer and judge run

    SınırlarBoundaries

    Ne yapmazWhat it never does

    Bu liste pazarlık konusu değil — kodda tek tek yazılı. This list is non-negotiable — each line is written into the skill itself.

    Dürüst sınırlarHonest limitations

    1. Güven eşikleri kalibre değil.The confidence thresholds are not calibrated. 0.95 ve 0.80 sabit başlangıç değerleri. %90 doğruluk hedefi geçmişe dönük bir denemede doğrulandı, ileriye dönük olarak henüz ölçülmedi — bunun için bir test koşucusu yazılması gerekiyor. 0.95 and 0.80 are fixed cold-start values. The ≥90% precision target was validated on a retrospective replay, not yet prospectively — that needs a test runner to be built.
    2. Doküman katmanı her gün ayakta olmuyor.The doc layer is not up every day. 30 Temmuz koşusunda akademi/Slack/Zendesk katmanı kapalıydı; o gün bazı "bug değil" gerekçeleri tek kaynağa dayandı. Bu rapor başlığına yazıldı, gizlenmedi. On the 30 July run the academy/Slack/Zendesk layer was down; that day some "not a bug" justifications rested on a single source. It was written into the report header, not hidden.
    3. Ürün mühendisi teyidi pratikte hiç çalışmıyor.Product-engineer confirmation effectively never fires. Bir yorumun ürün ekibinden gelip gelmediğini doğrulamak için takım listeleri gerekiyor — ama 20 takım dosyasının hiçbirinde üye listesi yok. Bu yüzden bu kanıt sınıfı yapısal olarak sessiz. Verifying that a comment came from a product team needs team rosters — but none of the 20 team files carries a member list. So that evidence class is structurally silent.
    4. Yorumu olmayan task'ta yorum tabanlı kanıt çalışmaz.Comment-based evidence does nothing on a comment-less ticket. Task ilk düştüğünde eskalasyon veya mühendis teyidi henüz yoktur; o an tek dayanak açıklamadaki çelişki, kod ve dokümandır. When a ticket first lands there is no escalation or engineer confirmation yet; at that moment the only footing is the contradiction in the description, plus code and docs.
    5. Bilinen ve kabul edilen kaçırmalar var.There are known, accepted misses. "Şuna sorman gerekebilir" gibi tavsiye cümleleri eskalasyon sayılmaz; Optimus'un kendi ekiplerinden gelen teyitler kanıt sayılmaz. İkisi de bilerek böyle — yanlış alarmı önlemek için. Advice-shaped sentences like "you may need to ask X" do not count as escalation, and confirmations from Optimus' own teams do not count as evidence. Both are deliberate — to prevent false alarms.