Four generations of test configuration beneath a customer's CI pipeline.ארבעה דורות של קונפיגורציית בדיקות, מתחת ל-CI pipeline של לקוח.
Every time an autonomous-vehicle team changes its driving software, it needs to know what that change broke. Driving real cars to find out is too slow and too risky, so the answer comes from simulation: thousands of test scenarios, launched automatically whenever a new build appears. That automation has two halves: the pipeline that decides when tests run, and the layer that defines what runs and how. At Foretellix the pipeline belonged to the customer. I designed the other half.
One challenge ran through four generations of the product: a launch is built from a predefined template, often one someone else created, yet every run needs something slightly different.
Before the GUI, a test run was a spreadsheet, with every parameter typed by hand.
The first GUI, in 2023, moved it into one modal: pick a runner file, paste the environment settings as raw JSON, run.
"It is too hard for user to setup a test using raw json file... New engineers have a hard time learning how to use the tool."
A customer's V&V organisation
Underneath the JSON sat a harder problem: no launch is quite the same twice. A new simulator image, another map, one changed parameter: without per-run overrides, changing one image tag meant cloning a whole settings record.
Access to customers was limited, but I could still build the picture from every source available, and ran my own validation where none existed:
The 2024 concept was paused and never built. But one reviewer's request, to decide what a launch may override, returned two years later as the override model.
Two rounds of the composition concept: a suite as a tree, then a library where a suite overrides a scenario's parameters.
Research only outlives its project if it becomes someone's requirement. Findings belong on the shared backlog, so a paused project still leaves evidence for the next one.
Four generations followed, from a raw-JSON modal (2023) to the curated launch drawer (2026). Three design decisions carried across them.
In the 2023 modal, the tests and their environment were one block of configuration: changing either meant editing both.
Split what to run from how to run it, as separate objects that only come together at launch.
What to run is a scenario: actors, map and parameters. How to run it is a preset: setup files, simulator image, timeouts and variables.
Four objects carry the product:
A clean split still needs explicit rules where the halves meet. Launching also saves the configuration, and that undocumented seam came back as a question a year later.
Executable assets got their own top-level tab, Test Management: a library of everything users can run, built on one repeating pattern:
I first proposed a creation wizard inside the launch, then dropped it: creating objects mid-launch would leave orphaned plans and suites.
One repeating pattern for every master, and one temporary surface for every run: the launch drawer.
Keep creating and launching apart. A surface that only customises a run can be reused from anywhere.
By 2026 a preset held far more than anyone needed: application engineers estimated that roughly 95% of environment settings were irrelevant to the person launching a test.
Administrators curate presets and declare which fields a user may override at launch. "Show all" keeps every setting within reach, because no one can predict what a specific run will need.
The same drawer twice: the curated default, and every setting one click away.
The template stays stable, and each execution changes only what it needs. Platform engineering calls this a golden path; I reached it from the field.
Simplify by hiding, never by removing. Experts only trust a curated default if the full set stays one click away.
In the last year my workflow changed too. I designed directly in a branch of the product: Figma for early exploration only, then the UX built in live code with mocked data, down to the shared components in Storybook.
That left one gap: development and QA still needed a source of truth, and redrawing finished code in Figma would double the work.
Replace the Figma and Zeplin handoff with QA reference pages pinned to a prototype commit: every state, how it was verified, and screenshots generated straight from the branch by a command I built.
Its first reader turned out to be me. Writing every state down as a test case worked as a second pass on my own design, and it surfaced edge cases I had missed. On the Manage Maps page alone, it exposed three states the design was asked for and never showed, such as a failed import that leaves no trace.
The front-end developer audited one against its commit and found four uncommitted files: the method doing exactly its job.
When the design lives in code, the reference has to come from the code too, or the designer maintains two versions of the truth.
By the time I left, new work plugged into the model without reshaping it:
Measure from day one. Nobody tracked usage, and I argued from that absence instead of fixing it. The first data I finally saw pointed the right way: people opened "Show all" again and again. Next time, the instrumentation ships with the design.
בכל פעם שצוות של רכב אוטונומי משנה את תוכנת הנהיגה, הוא צריך לדעת מה השינוי שבר. לנהוג ברכבים אמיתיים כדי לגלות זה איטי מדי ומסוכן מדי, ולכן התשובה מגיעה מסימולציה: אלפי תרחישי בדיקה, שמורצים אוטומטית בכל פעם שמופיע build חדש. לאוטומציה הזאת יש שני חצאים: ה-pipeline שמחליט מתי בדיקות רצות, והשכבה שמגדירה מה רץ ואיך. ב-Foretellix ה-pipeline היה של הלקוח. אני עיצבתי את החצי השני.
אתגר אחד ליווה ארבעה דורות של המוצר: כל הרצה נבנית מתבנית מוגדרת מראש, שלא פעם נוצרה על ידי מישהו אחר, ובכל זאת כל הרצה צריכה משהו קצת אחר.
לפני שהיה ממשק, הרצת בדיקות הייתה גיליון אלקטרוני, עם כל פרמטר מוקלד ידנית.
הממשק הראשון, ב-2023, העביר את זה למודאל אחד: בוחרים קובץ runner, מדביקים את הגדרות הסביבה כ-JSON גולמי, ומריצים.
"קשה מדי למשתמש להגדיר בדיקה באמצעות קובץ JSON גולמי... מהנדסים חדשים מתקשים ללמוד להשתמש בכלי."
ארגון ה-V&V של לקוח
מתחת ל-JSON הסתתרה בעיה קשה יותר: אין שתי הרצות זהות. image חדש לסימולטור, מפה אחרת, פרמטר אחד ששונה: בלי דריסות ברמת ההרצה, שינוי של tag יחיד ב-image חייב לשכפל רשומת הגדרות שלמה.
הגישה ללקוחות הייתה מוגבלת, ובכל זאת יכולתי לבנות את התמונה מכל מקור זמין, והרצתי ולידציה משלי איפה שלא הייתה:
הקונספט של 2024 הוקפא ולא נבנה. אבל בקשה של אחד הסוקרים, להחליט מה הרצה רשאית לדרוס, חזרה שנתיים אחר כך כמודל הדריסות.
שני סבבים של קונספט הקומפוזיציה: suite כעץ, ואחר כך ספרייה שבה suite דורס פרמטרים של תרחיש.
מחקר שורד את הפרויקט שלו רק אם הוא הופך לדרישה של מישהו. ממצאים צריכים לשבת ב-backlog המשותף, כדי שגם פרויקט מוקפא ישאיר ראיות לפרויקט הבא.
אחר כך באו ארבעה דורות, ממודאל עם JSON גולמי (2023) ועד מגירת ההרצה המותאמת (2026). שלוש החלטות עיצוב ליוו את כולם.
במודאל של 2023, הבדיקות והסביבה שלהן היו גוש אחד של קונפיגורציה: כדי לשנות אחד מהם היה צריך לערוך את שניהם.
להפריד בין מה להריץ לבין איך להריץ, כאובייקטים נפרדים שמתחברים רק בהרצה.
מה להריץ הוא תרחיש: שחקנים, מפה ופרמטרים. איך להריץ הוא פריסט: קובצי setup, image של הסימולטור, timeouts ומשתנים.
ארבעה אובייקטים נושאים את המוצר:
גם הפרדה נקייה צריכה כללים מפורשים בנקודה שבה החצאים נפגשים. הרצה שומרת גם את הקונפיגורציה, והתפר הזה, שלא תועד, חזר כשאלה שנה אחר כך.
לנכסים שאפשר להריץ ניתן טאב משלהם ברמה העליונה, Test Management: ספרייה של כל מה שאפשר להריץ, שבנויה על דפוס אחד שחוזר:
בהתחלה הצעתי wizard ליצירה בתוך ההרצה, ואז ויתרתי עליו: יצירת אובייקטים באמצע הרצה הייתה משאירה plans ו-suites יתומים.
תבנית אחת שחוזרת לכל master, ומשטח זמני אחד לכל הרצה: מגירת ההרצה.
להפריד בין יצירה להרצה. משטח שרק מתאים הרצה אפשר להשתמש בו שוב מכל מקום.
עד 2026 פריסט הכיל הרבה יותר ממה שמישהו צריך: מהנדסי האפליקציה העריכו שכ-95% מהגדרות הסביבה לא רלוונטיות למי שמריץ בדיקה.
אדמינים מגדירים פריסטים וקובעים אילו שדות משתמש רשאי לדרוס בהרצה. "Show all" משאיר כל הגדרה בהישג יד, כי אף אחד לא יכול לחזות מה הרצה ספציפית תצטרך.
אותה מגירה פעמיים: ברירת המחדל המותאמת, וכל ההגדרות במרחק קליק.
התבנית נשארת יציבה, וכל הרצה משנה רק את מה שהיא צריכה. ב-platform engineering קוראים לזה golden path; אני הגעתי אליו מהשטח.
לפשט על ידי הסתרה, אף פעם לא על ידי הסרה. מומחים סומכים על ברירת מחדל מותאמת רק אם הסט המלא נשאר במרחק קליק.
בשנה האחרונה השתנתה גם שיטת העבודה שלי. עיצבתי ישירות ב-branch של המוצר: Figma רק לחקירה מוקדמת, ואז ה-UX נבנה בקוד חי עם דאטה מדומה, עד לקומפוננטות המשותפות ב-Storybook.
נשאר פער אחד: פיתוח ו-QA עדיין היו צריכים מקור אמת, ולצייר מחדש ב-Figma קוד גמור היה מכפיל את העבודה.
להחליף את ה-handoff ב-Figma וב-Zeplin בעמודי QA reference שמוצמדים ל-commit של אב הטיפוס: כל מצב, איך הוא אומת, וצילומי מסך שנוצרים ישירות מה-branch באמצעות פקודה שבניתי.
הקורא הראשון של העמוד הזה היה אני. לכתוב כל מצב כמקרה בדיקה עבד כמעבר שני על העיצוב שלי, וחשף מקרי קצה שפספסתי. בעמוד Manage Maps לבדו הוא חשף שלושה מצבים שהעיצוב התבקש להראות ולא הראה, למשל ייבוא שנכשל ולא משאיר שום עקבות.
מפתח ה-front-end בדק אחד מהם מול ה-commit שלו ומצא ארבעה קבצים שלא נכנסו ל-commit: השיטה עושה בדיוק את מה שהיא אמורה לעשות.
כשהעיצוב חי בקוד, גם הרפרנס צריך לבוא מהקוד, אחרת המעצב מתחזק שתי גרסאות של האמת.
כשעזבתי, עבודה חדשה התחברה למודל בלי לעצב אותו מחדש:
למדוד מהיום הראשון. אף אחד לא עקב אחרי השימוש, ואני טענתי מתוך ההיעדר הזה במקום לתקן אותו. הדאטה הראשון שבסוף ראיתי הצביע לכיוון הנכון: אנשים פתחו את "Show all" שוב ושוב. בפעם הבאה, המדידה יוצאת יחד עם העיצוב.