This is an automated email from the ASF dual-hosted git repository.
voidmatcha pushed a commit to branch master
in repository https://gitbox.apache.org/repos/asf/zeppelin.git
The following commit(s) were added to refs/heads/master by this push:
new a46011e93e [ZEPPELIN-6653] Keep AGENTS.md chains within Codex budget
a46011e93e is described below
commit a46011e93e41d899a180644ee01fd226d3fc71bf
Author: chaeyoung kim <[email protected]>
AuthorDate: Mon Oct 5 22:09:04 2026 +0900
[ZEPPELIN-6653] Keep AGENTS.md chains within Codex budget
### What is this PR for?
Codex CLI reads the root and scoped `AGENTS.md` files as one cumulative
instruction chain. Some Zeppelin chains exceeded the default 32 KiB limit,
causing the most specific guidance to be truncated.
This PR moves detailed architecture and E2E guidance into linked reference
files under `.agents/`, while keeping essential instructions in the
automatically loaded files. The largest remaining chain is 29,026 bytes.
### What type of PR is it?
Documentation
### Todos
* [x] Keep every existing `AGENTS.md` chain under 32 KiB
### What is the Jira issue?
https://issues.apache.org/jira/browse/ZEPPELIN-6653
### How should this be tested?
This is a documentation-only change.
The combined instruction-chain sizes were verified as follows:
* Root: 8,820 bytes
* `docs`: 15,950 bytes
* `zeppelin-web-angular`: 22,764 bytes
* `zeppelin-web-angular/projects/zeppelin-react`: 26,578 bytes
* `zeppelin-web-angular/e2e`: 29,026 bytes
The following checks pass:
```bash
./mvnw clean org.apache.rat:apache-rat-plugin:check -Prat
git diff --check
```
### Screenshots (if appropriate)
Not applicable
### Questions:
- Does the license files need to update? No
- Is there breaking changes for older versions? No
- Does this needs documentation? No, this PR is the documentation update
Closes #5531 from chelsseeey/ZEPPELIN-6653-trim-agents-chain.
Signed-off-by: YONGJAE LEE <[email protected]>
---
.../e2e/AGENTS.md => .agents/e2e.md | 55 +---
.agents/module-architecture.md | 263 +++++++++++++++
.agents/server-interpreter-communication.md | 127 +++++++
AGENTS.md | 363 +--------------------
zeppelin-web-angular/e2e/AGENTS.md | 101 +-----
5 files changed, 421 insertions(+), 488 deletions(-)
diff --git a/zeppelin-web-angular/e2e/AGENTS.md b/.agents/e2e.md
similarity index 71%
copy from zeppelin-web-angular/e2e/AGENTS.md
copy to .agents/e2e.md
index 4fd63b1de7..2b6f925369 100644
--- a/zeppelin-web-angular/e2e/AGENTS.md
+++ b/.agents/e2e.md
@@ -15,27 +15,13 @@ See the License for the specific language governing
permissions and
limitations under the License.
-->
-# AGENTS.md
+# Detailed E2E Guidance
-> E2E (Playwright) conventions for `zeppelin-web-angular/e2e/`. Loaded only
when working under `e2e/`; the package baseline is
`zeppelin-web-angular/AGENTS.md`, which covers unit tests. See [AGENTS.md
specification](https://github.com/agentsmd/agents.md).
-
-Config: `zeppelin-web-angular/playwright.config.js` (Angular UI) and
`playwright.classic.config.js` (legacy classic UI), sharing
`playwright.shared.js`. This document is the source of truth for E2E
conventions, for contributors and for coding agents alike.
-
-## Layout
-
-- Specs: `e2e/tests/<area>/[<group>/]<feature>.spec.ts` (areas:
`authentication`, `home`, `login`, `notebook`, `share`, `theme`, `workspace`).
Larger areas group specs one level deeper, as in `notebook/keyboard/` and
`workspace/notebook-repos/`. `tests/app.spec.ts` covers the app shell and sits
outside any area.
-- Page Objects (POM), split by role:
- - `e2e/models/<name>.ts`: locators + primitive actions (click, fill,
navigate, simple state checks).
- - `e2e/models/<name>.util.ts`: workflows, composite verification, scenario
helpers.
- - Most existing POMs are a single file. Split a new one by role, and split
an existing one when its workflow code outgrows its locators.
-- Shared helpers: `e2e/utils.ts`.
-
-## Style
-
-- English only. No unnecessary comments.
-- BDD via `test.step('Given/When/Then …', …)`. Steps show up in traces and
reports; `// Given:` comments do not. Some specs still use comments; migrate a
test's comments to steps when you touch it.
-- One `test.describe` per feature; construct the feature's own POM in
`beforeEach`. A secondary POM that only one test needs, such as the second
viewer in a collaboration test or a page reached mid-test, can be built in the
test body.
-- `test.describe.serial` is a last resort: one failure skips every later test
in the group, which hides the rest instead of reporting them. Playwright
recommends against it (https://playwright.dev/docs/test-parallel#serial-mode).
Prefer making each test set up its own state.
+> Extended guidance for the Playwright suites. The scoped
+> [`AGENTS.md`](../zeppelin-web-angular/e2e/AGENTS.md) contains the rules that
+> always apply. Read the relevant sections here before using an escape hatch,
+> changing locator or assertion patterns, running a specialized suite, testing
+> an Angular-to-React migration seam, or modifying Classic UI coverage.
## Escape hatches
@@ -69,27 +55,6 @@ Much of the suite predates this section: it inlines CSS and
mostly omits `exact:
- A network wait is synchronization, not proof.
`waitForLoadState('networkidle')` is discouraged by Playwright and the suite
still has several, one of them inside `waitForZeppelinReady`; in new code wait
on a user-visible signal instead. When you do wait on the network, assert the
rendered result afterwards.
- The lint config covers part of this section, not all of it.
`eslint-plugin-playwright` has no rule for always-true assertions, so those are
a review responsibility.
-## Readiness & Auth
-
-- After navigation, wait with `waitForZeppelinReady(page)` from `e2e/utils.ts`
(not fixed sleeps).
-- Auth is programmatic: the `setup` project logs in once and writes
`playwright/.auth/user.json`; browser projects consume it via `storageState`.
Do not add per-test login races. For logged-out scenarios use a fresh context.
-- A skip says why it skipped. `playwright/no-skipped-test` errors on the
declaration forms (`test.skip('title', fn)`, `test.describe.skip`) and on a
bare `test.skip()` outside an `if`; those need the `eslint-disable` hatch and a
tracking key. Every other skip passes lint whatever its message says, so the
message is a convention, not a gate: name the missing capability (auth mode,
interpreter, environment feature) or the tracking key.
-
-## Coverage Annotation (Required)
-
-Every `describe` must declare the page/component it exercises so coverage is
attributed:
-
-```ts
-import { addPageAnnotationBeforeEach, PAGES } from '../../utils';
-
-test.describe('Home Page - Core Elements', () => {
- addPageAnnotationBeforeEach(PAGES.WORKSPACE.HOME);
- // …
-});
-```
-
-Use an existing key from the `PAGES` object in `e2e/utils.ts`; add a new one
there if the page is missing. The reporter discovers
`src/app/**/*.component.ts` automatically and removes only the entries in
`COVERAGE_EXCLUDED_COMPONENTS`, so component additions, deletions and moves
update the denominator automatically. `PAGES` separately supplies the
annotation names and must match those discovered targets.
`test/reporter.coverage.spec.ts` enforces that match, rejects duplicate entries
and [...]
-
## Running
- Node: `nvm use` (version pinned in `.nvmrc`).
@@ -112,14 +77,6 @@ Use an existing key from the `PAGES` object in
`e2e/utils.ts`; add a new one the
| `npm run e2e:codegen` | Record against `:4200` |
| `npm run e2e:cleanup` | Delete leftover test notebooks
(`e2e/cleanup-util.ts`) |
-## Adding a Test (Agents Start Here)
-
-1. Pick/confirm the target route and the `PAGES` key.
-2. Copy the shape of an existing spec in the same `<area>`; reuse or extend
the matching POM (`models/<name>.ts` + `.util.ts`). Do not inline selectors the
POM already owns.
-3. Annotate the page (`addPageAnnotationBeforeEach`), navigate, then
`waitForZeppelinReady`.
-4. If the test covers a scenario in `e2e/scenarios/notebook-parity.json`, add
its stable ID as a Playwright tag such as `{ tag: '@NB-PARITY-001' }`. Keep the
title human-readable; the registry links coverage by tag and path. Browser
execution controls such as project lists and skip conditions stay in the spec
rather than being copied into the registry. Keep browser assumptions in the
registry only when they define the scenario's behavior or expected outcome.
-5. Run `npm run e2e:fast` and iterate until green.
-
## Migration (Angular to React Microfrontend)
Pages are moving from Angular to React fragments incrementally. Today this is
narrow: the published paragraph route reads a `?react=true` flag
(`published/paragraph/paragraph.component`), the notebook footer swaps via a
`?reactFooter=true` flag (read into the notebook component's `useReactFooter`
input), the configuration table swaps via a `?reactConfiguration=true` flag
(`configuration/configuration.component`), and the notebook repository list
swaps via a `?reactNotebookRepos=true` fla [...]
diff --git a/.agents/module-architecture.md b/.agents/module-architecture.md
new file mode 100644
index 0000000000..727ac2dc39
--- /dev/null
+++ b/.agents/module-architecture.md
@@ -0,0 +1,263 @@
+<!--
+Licensed to the Apache Software Foundation (ASF) under one or more
+contributor license agreements. See the NOTICE file distributed with
+this work for additional information regarding copyright ownership.
+The ASF licenses this file to You under the Apache License, Version 2.0
+(the "License"); you may not use this file except in compliance with
+the License. You may obtain a copy of the License at
+
+ http://www.apache.org/licenses/LICENSE-2.0
+
+Unless required by applicable law or agreed to in writing, software
+distributed under the License is distributed on an "AS IS" BASIS,
+WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
+See the License for the specific language governing permissions and
+limitations under the License.
+-->
+
+# Module Architecture
+
+> Detailed architecture reference for coding agents. The repository-root
+> [`AGENTS.md`](../AGENTS.md) contains the instructions that always apply.
+
+### Dependency Flow
+
+```
+zeppelin-interpreter Base API: Interpreter, InterpreterContext,
Thrift services
+ ↓
+zeppelin-interpreter-shaded Uber JAR (maven-shade-plugin, relocated packages)
+ ↓
+zeppelin-server Core engine + Jetty 11, REST/WebSocket APIs, HK2
DI, entry point
+```
+
+### Core Modules
+
+#### `zeppelin-interpreter/`
+The base framework that all interpreters depend on. Defines the interpreter
API and the Thrift communication protocol. This module is shaded into an uber
JAR (`zeppelin-interpreter-shaded`) and placed on each interpreter process's
classpath.
+
+Key classes:
+- `Interpreter` (abstract) / `AbstractInterpreter` — base class every
interpreter extends
+- `InterpreterContext` — carries notebook/paragraph/user info into
`interpret()` calls
+- `InterpreterGroup` — manages a group of interpreter instances sharing one
process
+- `InterpreterResult` / `InterpreterOutput` — execution result model
+- `RemoteInterpreterServer` — **entry point of each interpreter JVM process**;
implements the Thrift `RemoteInterpreterService` server; receives RPC calls
from zeppelin-server
+- `InterpreterLauncher` (abstract) — how an interpreter process is started
(Standard, Docker, K8s, YARN)
+- `LifecycleManager` — manages interpreter process lifecycle (Null = keep
alive, Timeout = idle shutdown)
+- `DependencyResolver` / `AbstractDependencyResolver` — Maven artifact
resolution for `%dep` paragraphs
+
+Thrift definitions (`src/main/thrift/`):
+- `RemoteInterpreterService.thrift` — server → interpreter RPCs
+- `RemoteInterpreterEventService.thrift` — interpreter → server event callbacks
+
+#### `zeppelin-server/`
+The entry point and core of the Zeppelin application. Combines the web server
/ API layer with the core notebook engine, interpreter lifecycle management,
scheduling, search, and plugin loading.
+
+Web / API layer (`org.apache.zeppelin.server`, `rest`, `socket`):
+- `ZeppelinServer` — `main()`, embedded Jetty 11 server, HK2 DI setup
+- `NotebookRestApi`, `InterpreterRestApi`, `SecurityRestApi`,
`ConfigurationsRestApi` — REST endpoints in `org.apache.zeppelin.rest`
+- `NotebookServer` — WebSocket endpoint (`/ws`) for real-time notebook
operations and paragraph execution
+- `RemoteInterpreterEventServer` — Thrift server receiving callbacks from
interpreter processes (output streaming, status updates)
+
+Engine / runtime (`org.apache.zeppelin.notebook`, `interpreter`, `scheduler`,
`search`, `plugin`, `storage`, `conf`):
+- `Notebook` / `Note` / `Paragraph` — notebook data model and execution
+- `InterpreterFactory` — creates interpreter instances
+- `InterpreterSettingManager` — loads `interpreter-setting.json` from each
interpreter directory, manages interpreter configurations
+- `InterpreterSetting` — one interpreter's config + runtime state; creates
`InterpreterLauncher` and `RemoteInterpreterProcess`
+- `ManagedInterpreterGroup` — server-side `InterpreterGroup` implementation;
owns the `RemoteInterpreterProcess`
+- `NoteManager` — notebook CRUD, folder tree
+- `SchedulerService` — Quartz-based cron scheduling
+- `SearchService` — Lucene-based notebook search
+- `PluginManager` — loads launcher and notebook-repo plugins (custom
classloading, not Java SPI)
+- `ZeppelinConfiguration` — config management (env vars → system properties →
`zeppelin-site.xml` → defaults)
+- `RecoveryStorage` — persists interpreter process info for server-restart
recovery
+- `ConfigStorage` — persists interpreter settings to JSON
+
+#### `zeppelin-interpreter-shaded/`
+Uses maven-shade-plugin to package `zeppelin-interpreter` + dependencies into
an uber JAR with relocated packages (e.g., `org.apache.thrift` →
`org.apache.zeppelin.shaded.org.apache.thrift`). This JAR is placed on each
interpreter process's classpath.
+
+#### `zeppelin-client/`
+REST/WebSocket client library for programmatic access to Zeppelin.
+
+### Interpreter Modules
+
+Each interpreter is an independent Maven module inheriting from
`zeppelin-interpreter-parent`:
+
+| Module | Description |
+|--------|-------------|
+| `spark/` | Apache Spark (Scala/Python/R/SQL) — most complex interpreter |
+| `python/` | IPython/Python |
+| `flink/` | Apache Flink (Scala/Python/SQL) |
+| `jdbc/` | JDBC (PostgreSQL, MySQL, Hive, etc.) |
+| `shell/` | Bash/Shell commands |
+| `markdown/` | Markdown rendering (Flexmark) |
+| `java/` | Java interpreter |
+| `groovy/` | Groovy |
+| `neo4j/` | Neo4j Cypher |
+| `mongodb/` | MongoDB |
+| `elasticsearch/` | Elasticsearch |
+| `bigquery/` | Google BigQuery |
+| `cassandra/` | Apache Cassandra CQL |
+| `hbase/` | Apache HBase |
+| `livy/` | Apache Livy (remote Spark) |
+| `sparql/` | SPARQL queries |
+| `influxdb/` | InfluxDB |
+| `file/` | HDFS/local file browser |
+
+### Plugin Modules (`zeppelin-plugins/`)
+
+**Launcher plugins** (`launcher/`) — how interpreter processes are started:
+- `StandardInterpreterLauncher` (builtin) — local JVM process via
`bin/interpreter.sh`
+- `SparkInterpreterLauncher` (builtin) — Spark-specific launcher with
`spark-submit`
+- `DockerInterpreterLauncher` — Docker container
+- `K8sStandardInterpreterLauncher` — Kubernetes pod
+- `YarnInterpreterLauncher` — YARN container
+- `FlinkInterpreterLauncher` — Flink-specific
+- `ClusterInterpreterLauncher` — Zeppelin cluster mode
+
+**NotebookRepo plugins** (`notebookrepo/`) — where notebooks are persisted:
+- `VFSNotebookRepo` (builtin) — local filesystem (Apache VFS)
+- `GitNotebookRepo` (builtin) — local git repo
+- `GitHubNotebookRepo` — GitHub
+- `S3NotebookRepo` — Amazon S3
+- `GCSNotebookRepo` — Google Cloud Storage
+- `AzureNotebookRepo` — Azure Blob Storage
+- `MongoNotebookRepo` — MongoDB
+- `OSSNotebookRepo` — Alibaba Cloud OSS
+
+### Frontend
+
+- `zeppelin-web-angular/` — active frontend (Angular; versions in
`package.json`, Node build pin in `pom.xml` `node.version`)
+- `zeppelin-web/` — Legacy AngularJS (activated with `-Pweb-classic`)
+
+### Configuration Files
+
+| File | Purpose |
+|------|---------|
+| `conf/zeppelin-site.xml` | Main server config (port, SSL, notebook storage,
interpreter settings). Copy from `.template` |
+| `conf/zeppelin-env.sh` | Shell environment (JAVA_OPTS, memory, Spark
master). Copy from `.template` |
+| `conf/shiro.ini` | Authentication/authorization (users, roles, LDAP,
Kerberos, PAM). Copy from `.template` |
+| `conf/interpreter.json` | Runtime interpreter settings — **auto-generated**,
do not edit manually |
+| `conf/log4j2.properties` | Logging configuration |
+| `conf/interpreter-list` | Static list of available interpreters with Maven
coordinates |
+| `{interpreter}/resources/interpreter-setting.json` | Interpreter defaults
(build-time, bundled in JAR) |
+
+`conf/*.template` files are the source of truth. Actual config files
(`zeppelin-site.xml`, `shiro.ini`, etc.) are `.gitignored`.
+
+### Module Boundaries
+
+Where new code should go:
+
+| If the code... | Put it in |
+|----------------|-----------|
+| Is a base interface/class that all interpreters need |
`zeppelin-interpreter` |
+| Handles notebook state, interpreter lifecycle, scheduling, search,
REST/WebSocket, or authentication realm | `zeppelin-server` |
+| Is specific to one backend (Spark, Flink, JDBC, etc.) | That interpreter's
module |
+| Is a new way to launch interpreter processes | `zeppelin-plugins/launcher/` |
+| Is a new notebook storage backend | `zeppelin-plugins/notebookrepo/` |
+
+**Important**: Code added to `zeppelin-interpreter` is exposed to **every
interpreter process** via the shaded JAR. Only add code there if all
interpreters genuinely need it.
+
+## Plugin System & Reflection Patterns
+
+### PluginManager — Custom Classloading
+
+`PluginManager` (`zeppelin-server/.../plugin/PluginManager.java`) loads
plugins without Java SPI:
+
+```
+Plugin loading flow:
+1. Check builtin list (hardcoded class names):
+ - Launchers: StandardInterpreterLauncher, SparkInterpreterLauncher
+ - NotebookRepos: VFSNotebookRepo, GitNotebookRepo
+ → if builtin: Class.forName(className) — direct classloading
+
+2. If not builtin → external plugin:
+ → Scan pluginsDir/{Launcher|NotebookRepo}/{pluginName}/ for JARs
+ → Create URLClassLoader with those JARs
+ → classLoader.loadClass(className)
+ → Instantiate via reflection (constructor parameters)
+```
+
+External plugin directory structure:
+```
+plugins/
+ Launcher/
+ DockerInterpreterLauncher/
+ *.jar
+ K8sStandardInterpreterLauncher/
+ *.jar
+ NotebookRepo/
+ S3NotebookRepo/
+ *.jar
+ GCSNotebookRepo/
+ *.jar
+```
+
+### ReflectionUtils
+
+`ReflectionUtils` (`zeppelin-server/.../util/ReflectionUtils.java`) provides
generic reflection-based instantiation:
+
+```java
+// No-arg constructor
+ReflectionUtils.createClazzInstance(className)
+
+// Parameterized constructor
+ReflectionUtils.createClazzInstance(className, parameterTypes, parameters)
+```
+
+Used to instantiate:
+- `RecoveryStorage` — in `RemoteInterpreterServer` and
`InterpreterSettingManager`
+- `ConfigStorage` — in `InterpreterSettingManager`
+- `LifecycleManager` — in `RemoteInterpreterServer`
+- `NotebookRepo` — in `PluginManager`
+- `InterpreterLauncher` — in `PluginManager`
+
+### Interpreter Discovery
+
+`InterpreterSettingManager` discovers interpreters at startup:
+
+```
+1. Scan interpreterDir (default: interpreter/) for subdirectories
+2. For each subdirectory, look for interpreter-setting.json
+3. Parse JSON → List<RegisteredInterpreter>
+4. Register each interpreter's className, properties, editor settings
+```
+
+`interpreter-setting.json` format (in each interpreter module's resources):
+```json
+[{
+ "group": "spark",
+ "name": "spark",
+ "className": "org.apache.zeppelin.spark.SparkInterpreter",
+ "properties": {
+ "spark.master": { "defaultValue": "local[*]", "description": "Spark
master" }
+ },
+ "editor": { "language": "scala", "editOnDblClick": false }
+}]
+```
+
+### ZeppelinConfiguration Priority
+
+Configuration values are resolved in order (first match wins):
+1. **Environment variables** (e.g., `ZEPPELIN_HOME`, `ZEPPELIN_PORT`)
+2. **System properties** (e.g., `-Dzeppelin.server.port=8080`)
+3. **zeppelin-site.xml** (`conf/zeppelin-site.xml`)
+4. **Hardcoded defaults** (`ConfVars` enum in `ZeppelinConfiguration`)
+
+### HK2 Dependency Injection (zeppelin-server)
+
+`ZeppelinServer.startZeppelin()` sets up HK2 DI via
`ServiceLocatorUtilities.bind()`:
+
+```java
+new AbstractBinder() {
+ protected void configure() {
+ bind(storage).to(ConfigStorage.class);
+ bindAsContract(PluginManager.class).in(Singleton.class);
+ bindAsContract(InterpreterFactory.class).in(Singleton.class);
+
bindAsContract(NotebookRepoSync.class).to(NotebookRepo.class).in(Singleton.class);
+ bindAsContract(Notebook.class).in(Singleton.class);
+ // ... InterpreterSettingManager, SearchService, etc.
+ }
+}
+```
+
+REST API classes use `@Inject` to receive these singletons.
diff --git a/.agents/server-interpreter-communication.md
b/.agents/server-interpreter-communication.md
new file mode 100644
index 0000000000..09b2410209
--- /dev/null
+++ b/.agents/server-interpreter-communication.md
@@ -0,0 +1,127 @@
+<!--
+Licensed to the Apache Software Foundation (ASF) under one or more
+contributor license agreements. See the NOTICE file distributed with
+this work for additional information regarding copyright ownership.
+The ASF licenses this file to You under the Apache License, Version 2.0
+(the "License"); you may not use this file except in compliance with
+the License. You may obtain a copy of the License at
+
+ http://www.apache.org/licenses/LICENSE-2.0
+
+Unless required by applicable law or agreed to in writing, software
+distributed under the License is distributed on an "AS IS" BASIS,
+WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
+See the License for the specific language governing permissions and
+limitations under the License.
+-->
+
+# Server–Interpreter Communication
+
+> Detailed runtime reference for coding agents. The repository-root
+> [`AGENTS.md`](../AGENTS.md) contains the instructions that always apply.
+
+Zeppelin's most important architectural concept: the server and each
interpreter run in **separate JVM processes** communicating via **Apache Thrift
RPC**. This provides isolation, fault tolerance, and the ability to run
interpreters on remote hosts or containers.
+
+### Thrift Code Generation
+
+The `.thrift` files are in `zeppelin-interpreter/src/main/thrift/`. Generated
Java files are **checked into git** (not generated at build time) in
`zeppelin-interpreter/src/main/java/org/apache/zeppelin/interpreter/thrift/`.
+
+To regenerate after modifying `.thrift` files:
+```bash
+cd zeppelin-interpreter/src/main/thrift
+./genthrift.sh # requires 'thrift' compiler (v0.13.0) installed locally
+```
+
+The script runs the Thrift compiler, prepends ASF license headers, and moves
files to the source tree. **Never edit the generated Java files directly** —
changes will be lost on next regeneration.
+
+### Thrift IPC — Bidirectional
+
+**Server → Interpreter** (`RemoteInterpreterService`):
+```
+init(properties) — initialize interpreter process with
config
+createInterpreter(className, ...) — instantiate an interpreter class
+open(sessionId, className) — open/initialize an interpreter
+interpret(sessionId, className, code, context) — execute code (core method)
+cancel(sessionId, className, ...) — cancel running execution
+getProgress(sessionId, className) — poll execution progress (0-100)
+completion(sessionId, className, buf, cursor) — code completion
+close(sessionId, className) — close an interpreter
+shutdown() — terminate the interpreter process
+```
+
+**Interpreter → Server** (`RemoteInterpreterEventService`):
+```
+registerInterpreterProcess(info) — register after process startup
+appendOutput(event) — stream execution output incrementally
+updateOutput(event) — replace output content
+sendParagraphInfo(info) — update paragraph metadata
+updateAppStatus(event) — Zeppelin Application status
+runParagraphs(request) — trigger paragraph execution from
interpreter
+getResource(resourceId) — access ResourcePool shared state
+getParagraphList(noteId) — query notebook structure
+```
+
+### Paragraph Execution Chain
+
+When a user runs a paragraph, the full call chain is:
+
+```
+User clicks "Run" in browser
+ → WebSocket message to NotebookServer
+ → NotebookServer.runParagraph()
+ → Notebook.run()
+ → Paragraph.execute()
+ → RemoteInterpreter.interpret(code, context)
+ → RemoteInterpreterProcess.callRemoteFunction()
+ → [Thrift RPC over TCP]
+ → RemoteInterpreterServer.interpret()
+ → actual Interpreter.interpret() (e.g. SparkInterpreter)
+ → result returned via Thrift
+ → meanwhile: interpreter calls appendOutput() to stream partial
results back
+```
+
+### Interpreter Launch Chain
+
+When an interpreter process needs to be started:
+
+```
+RemoteInterpreter.interpret() [first call triggers launch]
+ → ManagedInterpreterGroup.getOrCreateInterpreterProcess()
+ → InterpreterSetting.createInterpreterProcess()
+ → InterpreterSetting.createLauncher(properties)
+ → PluginManager.loadInterpreterLauncher(launcherPlugin)
+ → [builtin: Class.forName() / external: URLClassLoader]
+ → InterpreterLauncher.launch(context)
+ → new ExecRemoteInterpreterProcess(...)
+ → ExecRemoteInterpreterProcess.start()
+ → ProcessBuilder → "bin/interpreter.sh"
+ → java -cp ... RemoteInterpreterServer [new JVM]
+ → RemoteInterpreterServer.main()
+ → registerInterpreterProcess() callback to server
+```
+
+### Interpreter Process Lifecycle
+
+1. **Launch**: Server creates `RemoteInterpreterProcess` via launcher plugin
+2. **Start**: Process starts as separate JVM (`bin/interpreter.sh` →
`RemoteInterpreterServer.main()`)
+3. **Register**: Process calls `registerInterpreterProcess()` back to server's
`RemoteInterpreterEventServer`
+4. **Init**: Server calls `init(properties)` — passes all configuration as a
flat `Map<String, String>`
+5. **Create**: Server calls `createInterpreter(className, properties)` —
instantiates interpreter via reflection
+6. **Open**: First `interpret()` triggers `LazyOpenInterpreter.open()` —
interpreter initializes resources
+7. **Execute**: `interpret(code, context)` — runs code; partial output streams
via `appendOutput()` events
+8. **Shutdown**: `close()` → `shutdown()` → JVM exits
+9. **Recovery**: `RecoveryStorage` persists process info; on server restart,
reconnects to surviving processes
+
+### InterpreterGroup Scoping
+
+`InterpreterOption` controls process isolation via `perNote` and `perUser`
settings:
+
+| perNote | perUser | Behavior |
+|---------|---------|----------|
+| `shared` | `shared` | All users share one process (default) |
+| `scoped` | `shared` | Separate interpreter instance per note, same process |
+| `isolated` | `shared` | Separate process per note |
+| `shared` | `scoped` | Separate interpreter instance per user, same process |
+| `shared` | `isolated` | Separate process per user |
+| `scoped` | `scoped` | Separate instance per user+note |
+| `isolated` | `isolated` | Separate process per user+note (full isolation) |
diff --git a/AGENTS.md b/AGENTS.md
index 090dc4433e..b4c658fc14 100644
--- a/AGENTS.md
+++ b/AGENTS.md
@@ -83,357 +83,18 @@ All interpreter modules build after
`zeppelin-interpreter-shaded`. A second shad
zeppelin-jupyter-interpreter → zeppelin-jupyter-interpreter-shaded → python
```
-## Module Architecture
-
-### Dependency Flow
-
-```
-zeppelin-interpreter Base API: Interpreter, InterpreterContext,
Thrift services
- ↓
-zeppelin-interpreter-shaded Uber JAR (maven-shade-plugin, relocated packages)
- ↓
-zeppelin-server Core engine + Jetty 11, REST/WebSocket APIs, HK2
DI, entry point
-```
-
-### Core Modules
-
-#### `zeppelin-interpreter/`
-The base framework that all interpreters depend on. Defines the interpreter
API and the Thrift communication protocol. This module is shaded into an uber
JAR (`zeppelin-interpreter-shaded`) and placed on each interpreter process's
classpath.
-
-Key classes:
-- `Interpreter` (abstract) / `AbstractInterpreter` — base class every
interpreter extends
-- `InterpreterContext` — carries notebook/paragraph/user info into
`interpret()` calls
-- `InterpreterGroup` — manages a group of interpreter instances sharing one
process
-- `InterpreterResult` / `InterpreterOutput` — execution result model
-- `RemoteInterpreterServer` — **entry point of each interpreter JVM process**;
implements the Thrift `RemoteInterpreterService` server; receives RPC calls
from zeppelin-server
-- `InterpreterLauncher` (abstract) — how an interpreter process is started
(Standard, Docker, K8s, YARN)
-- `LifecycleManager` — manages interpreter process lifecycle (Null = keep
alive, Timeout = idle shutdown)
-- `DependencyResolver` / `AbstractDependencyResolver` — Maven artifact
resolution for `%dep` paragraphs
-
-Thrift definitions (`src/main/thrift/`):
-- `RemoteInterpreterService.thrift` — server → interpreter RPCs
-- `RemoteInterpreterEventService.thrift` — interpreter → server event callbacks
-
-#### `zeppelin-server/`
-The entry point and core of the Zeppelin application. Combines the web server
/ API layer with the core notebook engine, interpreter lifecycle management,
scheduling, search, and plugin loading.
-
-Web / API layer (`org.apache.zeppelin.server`, `rest`, `socket`):
-- `ZeppelinServer` — `main()`, embedded Jetty 11 server, HK2 DI setup
-- `NotebookRestApi`, `InterpreterRestApi`, `SecurityRestApi`,
`ConfigurationsRestApi` — REST endpoints in `org.apache.zeppelin.rest`
-- `NotebookServer` — WebSocket endpoint (`/ws`) for real-time notebook
operations and paragraph execution
-- `RemoteInterpreterEventServer` — Thrift server receiving callbacks from
interpreter processes (output streaming, status updates)
-
-Engine / runtime (`org.apache.zeppelin.notebook`, `interpreter`, `scheduler`,
`search`, `plugin`, `storage`, `conf`):
-- `Notebook` / `Note` / `Paragraph` — notebook data model and execution
-- `InterpreterFactory` — creates interpreter instances
-- `InterpreterSettingManager` — loads `interpreter-setting.json` from each
interpreter directory, manages interpreter configurations
-- `InterpreterSetting` — one interpreter's config + runtime state; creates
`InterpreterLauncher` and `RemoteInterpreterProcess`
-- `ManagedInterpreterGroup` — server-side `InterpreterGroup` implementation;
owns the `RemoteInterpreterProcess`
-- `NoteManager` — notebook CRUD, folder tree
-- `SchedulerService` — Quartz-based cron scheduling
-- `SearchService` — Lucene-based notebook search
-- `PluginManager` — loads launcher and notebook-repo plugins (custom
classloading, not Java SPI)
-- `ZeppelinConfiguration` — config management (env vars → system properties →
`zeppelin-site.xml` → defaults)
-- `RecoveryStorage` — persists interpreter process info for server-restart
recovery
-- `ConfigStorage` — persists interpreter settings to JSON
-
-#### `zeppelin-interpreter-shaded/`
-Uses maven-shade-plugin to package `zeppelin-interpreter` + dependencies into
an uber JAR with relocated packages (e.g., `org.apache.thrift` →
`org.apache.zeppelin.shaded.org.apache.thrift`). This JAR is placed on each
interpreter process's classpath.
-
-#### `zeppelin-client/`
-REST/WebSocket client library for programmatic access to Zeppelin.
-
-### Interpreter Modules
-
-Each interpreter is an independent Maven module inheriting from
`zeppelin-interpreter-parent`:
-
-| Module | Description |
-|--------|-------------|
-| `spark/` | Apache Spark (Scala/Python/R/SQL) — most complex interpreter |
-| `python/` | IPython/Python |
-| `flink/` | Apache Flink (Scala/Python/SQL) |
-| `jdbc/` | JDBC (PostgreSQL, MySQL, Hive, etc.) |
-| `shell/` | Bash/Shell commands |
-| `markdown/` | Markdown rendering (Flexmark) |
-| `java/` | Java interpreter |
-| `groovy/` | Groovy |
-| `neo4j/` | Neo4j Cypher |
-| `mongodb/` | MongoDB |
-| `elasticsearch/` | Elasticsearch |
-| `bigquery/` | Google BigQuery |
-| `cassandra/` | Apache Cassandra CQL |
-| `hbase/` | Apache HBase |
-| `livy/` | Apache Livy (remote Spark) |
-| `sparql/` | SPARQL queries |
-| `influxdb/` | InfluxDB |
-| `file/` | HDFS/local file browser |
-
-### Plugin Modules (`zeppelin-plugins/`)
-
-**Launcher plugins** (`launcher/`) — how interpreter processes are started:
-- `StandardInterpreterLauncher` (builtin) — local JVM process via
`bin/interpreter.sh`
-- `SparkInterpreterLauncher` (builtin) — Spark-specific launcher with
`spark-submit`
-- `DockerInterpreterLauncher` — Docker container
-- `K8sStandardInterpreterLauncher` — Kubernetes pod
-- `YarnInterpreterLauncher` — YARN container
-- `FlinkInterpreterLauncher` — Flink-specific
-- `ClusterInterpreterLauncher` — Zeppelin cluster mode
-
-**NotebookRepo plugins** (`notebookrepo/`) — where notebooks are persisted:
-- `VFSNotebookRepo` (builtin) — local filesystem (Apache VFS)
-- `GitNotebookRepo` (builtin) — local git repo
-- `GitHubNotebookRepo` — GitHub
-- `S3NotebookRepo` — Amazon S3
-- `GCSNotebookRepo` — Google Cloud Storage
-- `AzureNotebookRepo` — Azure Blob Storage
-- `MongoNotebookRepo` — MongoDB
-- `OSSNotebookRepo` — Alibaba Cloud OSS
-
-### Frontend
-
-- `zeppelin-web-angular/` — active frontend (Angular; versions in
`package.json`, Node build pin in `pom.xml` `node.version`)
-- `zeppelin-web/` — Legacy AngularJS (activated with `-Pweb-classic`)
-
-### Configuration Files
-
-| File | Purpose |
-|------|---------|
-| `conf/zeppelin-site.xml` | Main server config (port, SSL, notebook storage,
interpreter settings). Copy from `.template` |
-| `conf/zeppelin-env.sh` | Shell environment (JAVA_OPTS, memory, Spark
master). Copy from `.template` |
-| `conf/shiro.ini` | Authentication/authorization (users, roles, LDAP,
Kerberos, PAM). Copy from `.template` |
-| `conf/interpreter.json` | Runtime interpreter settings — **auto-generated**,
do not edit manually |
-| `conf/log4j2.properties` | Logging configuration |
-| `conf/interpreter-list` | Static list of available interpreters with Maven
coordinates |
-| `{interpreter}/resources/interpreter-setting.json` | Interpreter defaults
(build-time, bundled in JAR) |
-
-`conf/*.template` files are the source of truth. Actual config files
(`zeppelin-site.xml`, `shiro.ini`, etc.) are `.gitignored`.
-
-### Module Boundaries
-
-Where new code should go:
-
-| If the code... | Put it in |
-|----------------|-----------|
-| Is a base interface/class that all interpreters need |
`zeppelin-interpreter` |
-| Handles notebook state, interpreter lifecycle, scheduling, search,
REST/WebSocket, or authentication realm | `zeppelin-server` |
-| Is specific to one backend (Spark, Flink, JDBC, etc.) | That interpreter's
module |
-| Is a new way to launch interpreter processes | `zeppelin-plugins/launcher/` |
-| Is a new notebook storage backend | `zeppelin-plugins/notebookrepo/` |
-
-**Important**: Code added to `zeppelin-interpreter` is exposed to **every
interpreter process** via the shaded JAR. Only add code there if all
interpreters genuinely need it.
-
-## Server–Interpreter Communication
-
-Zeppelin's most important architectural concept: the server and each
interpreter run in **separate JVM processes** communicating via **Apache Thrift
RPC**. This provides isolation, fault tolerance, and the ability to run
interpreters on remote hosts or containers.
-
-### Thrift Code Generation
-
-The `.thrift` files are in `zeppelin-interpreter/src/main/thrift/`. Generated
Java files are **checked into git** (not generated at build time) in
`zeppelin-interpreter/src/main/java/org/apache/zeppelin/interpreter/thrift/`.
-
-To regenerate after modifying `.thrift` files:
-```bash
-cd zeppelin-interpreter/src/main/thrift
-./genthrift.sh # requires 'thrift' compiler (v0.13.0) installed locally
-```
-
-The script runs the Thrift compiler, prepends ASF license headers, and moves
files to the source tree. **Never edit the generated Java files directly** —
changes will be lost on next regeneration.
-
-### Thrift IPC — Bidirectional
-
-**Server → Interpreter** (`RemoteInterpreterService`):
-```
-init(properties) — initialize interpreter process with
config
-createInterpreter(className, ...) — instantiate an interpreter class
-open(sessionId, className) — open/initialize an interpreter
-interpret(sessionId, className, code, context) — execute code (core method)
-cancel(sessionId, className, ...) — cancel running execution
-getProgress(sessionId, className) — poll execution progress (0-100)
-completion(sessionId, className, buf, cursor) — code completion
-close(sessionId, className) — close an interpreter
-shutdown() — terminate the interpreter process
-```
-
-**Interpreter → Server** (`RemoteInterpreterEventService`):
-```
-registerInterpreterProcess(info) — register after process startup
-appendOutput(event) — stream execution output incrementally
-updateOutput(event) — replace output content
-sendParagraphInfo(info) — update paragraph metadata
-updateAppStatus(event) — Zeppelin Application status
-runParagraphs(request) — trigger paragraph execution from
interpreter
-getResource(resourceId) — access ResourcePool shared state
-getParagraphList(noteId) — query notebook structure
-```
-
-### Paragraph Execution Chain
-
-When a user runs a paragraph, the full call chain is:
-
-```
-User clicks "Run" in browser
- → WebSocket message to NotebookServer
- → NotebookServer.runParagraph()
- → Notebook.run()
- → Paragraph.execute()
- → RemoteInterpreter.interpret(code, context)
- → RemoteInterpreterProcess.callRemoteFunction()
- → [Thrift RPC over TCP]
- → RemoteInterpreterServer.interpret()
- → actual Interpreter.interpret() (e.g. SparkInterpreter)
- → result returned via Thrift
- → meanwhile: interpreter calls appendOutput() to stream partial
results back
-```
-
-### Interpreter Launch Chain
-
-When an interpreter process needs to be started:
-
-```
-RemoteInterpreter.interpret() [first call triggers launch]
- → ManagedInterpreterGroup.getOrCreateInterpreterProcess()
- → InterpreterSetting.createInterpreterProcess()
- → InterpreterSetting.createLauncher(properties)
- → PluginManager.loadInterpreterLauncher(launcherPlugin)
- → [builtin: Class.forName() / external: URLClassLoader]
- → InterpreterLauncher.launch(context)
- → new ExecRemoteInterpreterProcess(...)
- → ExecRemoteInterpreterProcess.start()
- → ProcessBuilder → "bin/interpreter.sh"
- → java -cp ... RemoteInterpreterServer [new JVM]
- → RemoteInterpreterServer.main()
- → registerInterpreterProcess() callback to server
-```
-
-### Interpreter Process Lifecycle
-
-1. **Launch**: Server creates `RemoteInterpreterProcess` via launcher plugin
-2. **Start**: Process starts as separate JVM (`bin/interpreter.sh` →
`RemoteInterpreterServer.main()`)
-3. **Register**: Process calls `registerInterpreterProcess()` back to server's
`RemoteInterpreterEventServer`
-4. **Init**: Server calls `init(properties)` — passes all configuration as a
flat `Map<String, String>`
-5. **Create**: Server calls `createInterpreter(className, properties)` —
instantiates interpreter via reflection
-6. **Open**: First `interpret()` triggers `LazyOpenInterpreter.open()` —
interpreter initializes resources
-7. **Execute**: `interpret(code, context)` — runs code; partial output streams
via `appendOutput()` events
-8. **Shutdown**: `close()` → `shutdown()` → JVM exits
-9. **Recovery**: `RecoveryStorage` persists process info; on server restart,
reconnects to surviving processes
-
-### InterpreterGroup Scoping
-
-`InterpreterOption` controls process isolation via `perNote` and `perUser`
settings:
-
-| perNote | perUser | Behavior |
-|---------|---------|----------|
-| `shared` | `shared` | All users share one process (default) |
-| `scoped` | `shared` | Separate interpreter instance per note, same process |
-| `isolated` | `shared` | Separate process per note |
-| `shared` | `scoped` | Separate interpreter instance per user, same process |
-| `shared` | `isolated` | Separate process per user |
-| `scoped` | `scoped` | Separate instance per user+note |
-| `isolated` | `isolated` | Separate process per user+note (full isolation) |
-
-## Plugin System & Reflection Patterns
-
-### PluginManager — Custom Classloading
-
-`PluginManager` (`zeppelin-server/.../plugin/PluginManager.java`) loads
plugins without Java SPI:
-
-```
-Plugin loading flow:
-1. Check builtin list (hardcoded class names):
- - Launchers: StandardInterpreterLauncher, SparkInterpreterLauncher
- - NotebookRepos: VFSNotebookRepo, GitNotebookRepo
- → if builtin: Class.forName(className) — direct classloading
-
-2. If not builtin → external plugin:
- → Scan pluginsDir/{Launcher|NotebookRepo}/{pluginName}/ for JARs
- → Create URLClassLoader with those JARs
- → classLoader.loadClass(className)
- → Instantiate via reflection (constructor parameters)
-```
-
-External plugin directory structure:
-```
-plugins/
- Launcher/
- DockerInterpreterLauncher/
- *.jar
- K8sStandardInterpreterLauncher/
- *.jar
- NotebookRepo/
- S3NotebookRepo/
- *.jar
- GCSNotebookRepo/
- *.jar
-```
-
-### ReflectionUtils
-
-`ReflectionUtils` (`zeppelin-server/.../util/ReflectionUtils.java`) provides
generic reflection-based instantiation:
-
-```java
-// No-arg constructor
-ReflectionUtils.createClazzInstance(className)
-
-// Parameterized constructor
-ReflectionUtils.createClazzInstance(className, parameterTypes, parameters)
-```
-
-Used to instantiate:
-- `RecoveryStorage` — in `RemoteInterpreterServer` and
`InterpreterSettingManager`
-- `ConfigStorage` — in `InterpreterSettingManager`
-- `LifecycleManager` — in `RemoteInterpreterServer`
-- `NotebookRepo` — in `PluginManager`
-- `InterpreterLauncher` — in `PluginManager`
-
-### Interpreter Discovery
-
-`InterpreterSettingManager` discovers interpreters at startup:
-
-```
-1. Scan interpreterDir (default: interpreter/) for subdirectories
-2. For each subdirectory, look for interpreter-setting.json
-3. Parse JSON → List<RegisteredInterpreter>
-4. Register each interpreter's className, properties, editor settings
-```
-
-`interpreter-setting.json` format (in each interpreter module's resources):
-```json
-[{
- "group": "spark",
- "name": "spark",
- "className": "org.apache.zeppelin.spark.SparkInterpreter",
- "properties": {
- "spark.master": { "defaultValue": "local[*]", "description": "Spark
master" }
- },
- "editor": { "language": "scala", "editOnDblClick": false }
-}]
-```
-
-### ZeppelinConfiguration Priority
-
-Configuration values are resolved in order (first match wins):
-1. **Environment variables** (e.g., `ZEPPELIN_HOME`, `ZEPPELIN_PORT`)
-2. **System properties** (e.g., `-Dzeppelin.server.port=8080`)
-3. **zeppelin-site.xml** (`conf/zeppelin-site.xml`)
-4. **Hardcoded defaults** (`ConfVars` enum in `ZeppelinConfiguration`)
-
-### HK2 Dependency Injection (zeppelin-server)
-
-`ZeppelinServer.startZeppelin()` sets up HK2 DI via
`ServiceLocatorUtilities.bind()`:
-
-```java
-new AbstractBinder() {
- protected void configure() {
- bind(storage).to(ConfigStorage.class);
- bindAsContract(PluginManager.class).in(Singleton.class);
- bindAsContract(InterpreterFactory.class).in(Singleton.class);
-
bindAsContract(NotebookRepoSync.class).to(NotebookRepo.class).in(Singleton.class);
- bindAsContract(Notebook.class).in(Singleton.class);
- // ... InterpreterSettingManager, SearchService, etc.
- }
-}
-```
-
-REST API classes use `@Inject` to receive these singletons.
+## Architecture References
+
+The detailed architecture is kept outside this automatically loaded file so the
+scoped `AGENTS.md` chain remains within agent instruction budgets.
+
+- Read [Module Architecture](.agents/module-architecture.md) before changing
+ module boundaries, shared interpreter APIs, frontend placement, plugins, or
+ configuration ownership.
+- Read
+ [Server–Interpreter
Communication](.agents/server-interpreter-communication.md)
+ before changing Thrift contracts, interpreter launch or lifecycle behavior,
+ paragraph execution, or interpreter scoping.
## Contributing Guide
diff --git a/zeppelin-web-angular/e2e/AGENTS.md
b/zeppelin-web-angular/e2e/AGENTS.md
index 4fd63b1de7..d8625a79cf 100644
--- a/zeppelin-web-angular/e2e/AGENTS.md
+++ b/zeppelin-web-angular/e2e/AGENTS.md
@@ -37,37 +37,22 @@ Config: `zeppelin-web-angular/playwright.config.js`
(Angular UI) and `playwright
- One `test.describe` per feature; construct the feature's own POM in
`beforeEach`. A secondary POM that only one test needs, such as the second
viewer in a collaboration test or a page reached mid-test, can be built in the
test body.
- `test.describe.serial` is a last resort: one failure skips every later test
in the group, which hides the rest instead of reporting them. Playwright
recommends against it (https://playwright.dev/docs/test-parallel#serial-mode).
Prefer making each test set up its own state.
-## Escape hatches
+## Detailed Guidance
-Two comment forms mark a deliberate rule violation, and both require a reason:
+Before changing locators, assertions, deliberate rule exceptions, migration
+feature-flag coverage, or Classic UI tests, read the
+[detailed E2E guidance](../../.agents/e2e.md).
-- `// JUSTIFIED: <why>` for the conventions in this document, either trailing
the offending line or in the comment block directly above it. It is a contract
with the reviewer: the marker says the deviation is deliberate and the reason
says why. A `test.describe.serial` group needs one too.
-- `// eslint-disable-next-line <rule> -- <why>` for a lint rule. Give a reason
after the `--`; if the violation is tracked elsewhere, the ticket key is that
reason.
+The rules that apply to every test are:
-When a rule is both a convention here and a lint rule, `// JUSTIFIED:` is the
one to use: an `eslint-disable` silences the linter but leaves the convention
unmet.
-
-Neither hatch is a way to opt out of thinking. One without a concrete reason
will be challenged in review.
-
-## Locators
-
-Prefer user-facing, in this order:
-
-1. `getByRole('button' | 'link' | 'textbox', { name })`, `getByLabel`,
`getByText`. Pass `exact: true` alongside `name`. The default matches the
accessible name as a case-insensitive substring, which collides with note
titles and other page content and fails strict mode.
-2. `data-testid` (attribute selector) when a role or label is unavailable.
Adding one to the Angular or React template is allowed, and is better than
reaching into component internals with a CSS selector.
-3. A CSS selector only when the element offers neither, which is common for
ng-zorro internals and icon-only controls. It belongs in the Page Object, named
for what it does (`cancelButton`, not `.cancel-para`), never inline in a spec.
An accessible name that is really an icon glyph (`pause-circle`) is not an
improvement; it rots on the next icon swap.
-
-XPath is forbidden outright.
-
-Much of the suite predates this section: it inlines CSS and mostly omits
`exact: true`. The ratchet is that new or modified code complies. When you
touch a test that inlines a selector, move it into the Page Object as part of
that change.
-
-## Assertions
-
-- Web-first, auto-waiting assertions only: `toBeVisible`, `toHaveURL`,
`toHaveText`, `toHaveCount`. The suite still asserts on values it extracted
first in a few places, which `playwright/prefer-web-first-assertions` reports;
the ratchet applies here too.
-- No `waitForTimeout` without a `// JUSTIFIED:` rationale. When waiting on a
count, use `toHaveCount`.
-- No one-shot boolean checks (`expect(await el.isVisible())`) and no
always-true assertions on a locator (`toBeDefined`, `not.toBeNull`). A Locator
is always a defined, non-null object, so those pass whether or not the element
exists. Asserting a non-locator value is not this smell:
`expect.poll(...).not.toBeNull()` and a null-guard on a regex match are both
legitimate.
-- A conditional may gate a setup action on dual-mode UI (auth vs anonymous,
the optional welcome modal), which is why `playwright/no-conditional-in-test`
is off. Do not put an `expect` inside one: an assertion that runs on only one
branch passes by skipping the check it exists to make.
`playwright/no-conditional-expect` reports those and the suite still carries
some, so the ratchet applies here too.
-- A network wait is synchronization, not proof.
`waitForLoadState('networkidle')` is discouraged by Playwright and the suite
still has several, one of them inside `waitForZeppelinReady`; in new code wait
on a user-visible signal instead. When you do wait on the network, assert the
rendered result afterwards.
-- The lint config covers part of this section, not all of it.
`eslint-plugin-playwright` has no rule for always-true assertions, so those are
a review responsibility.
+- Prefer role, label, or text locators with exact accessible names, then
+ `data-testid`; keep CSS in Page Objects and never use XPath.
+- Use web-first assertions and observable readiness signals. Fixed waits and
+ one-shot visibility checks require a documented exception.
+- Mark a convention exception with `// JUSTIFIED: <reason>`; use a reasoned
+ `eslint-disable-next-line` only for lint-only exceptions.
+- Keep migration specs framework-neutral. The detailed guide records the
+ feature flags and the explicit exceptions for the frozen Classic UI.
## Readiness & Auth
@@ -90,28 +75,6 @@ test.describe('Home Page - Core Elements', () => {
Use an existing key from the `PAGES` object in `e2e/utils.ts`; add a new one
there if the page is missing. The reporter discovers
`src/app/**/*.component.ts` automatically and removes only the entries in
`COVERAGE_EXCLUDED_COMPONENTS`, so component additions, deletions and moves
update the denominator automatically. `PAGES` separately supplies the
annotation names and must match those discovered targets.
`test/reporter.coverage.spec.ts` enforces that match, rejects duplicate entries
and [...]
-## Running
-
-- Node: `nvm use` (version pinned in `.nvmrc`).
-- Dev server: `npm run start` at `http://localhost:4200` (Playwright reuses a
running one via `webServer.reuseExistingServer`).
-
-| Command | Purpose |
-| --- | --- |
-| `npm run e2e` | Full suite |
-| `npm run e2e:fast` | Chromium only (fast) |
-| `npm run e2e:fast -- tests/<area>/<feature>.spec.ts` | One spec (path is
relative to `e2e/`) |
-| `npm run e2e:fast -- -g '<test title>'` | One test, matched by title |
-| `npx eslint e2e/tests/<area>/<feature>.spec.ts` | Lint one file; `npm run
lint` covers the whole app |
-| `npm run e2e:classic` | Classic `/classic` UI suite against `:8080` (needs
`-Pweb-classic`) |
-| `npm run e2e:ui` | Playwright Test UI |
-| `npm run e2e:headed` | Headed run |
-| `npm run e2e:debug` | Step-by-step debugger |
-| `npm run e2e:report` | Open last HTML report |
-| `npm run e2e:report:classic` | Open last classic HTML report |
-| `npm run e2e:ci` | CI mode (`CI=true`, baseURL `:8080`), main then classic
suite |
-| `npm run e2e:codegen` | Record against `:4200` |
-| `npm run e2e:cleanup` | Delete leftover test notebooks
(`e2e/cleanup-util.ts`) |
-
## Adding a Test (Agents Start Here)
1. Pick/confirm the target route and the `PAGES` key.
@@ -119,41 +82,3 @@ Use an existing key from the `PAGES` object in
`e2e/utils.ts`; add a new one the
3. Annotate the page (`addPageAnnotationBeforeEach`), navigate, then
`waitForZeppelinReady`.
4. If the test covers a scenario in `e2e/scenarios/notebook-parity.json`, add
its stable ID as a Playwright tag such as `{ tag: '@NB-PARITY-001' }`. Keep the
title human-readable; the registry links coverage by tag and path. Browser
execution controls such as project lists and skip conditions stay in the spec
rather than being copied into the registry. Keep browser assumptions in the
registry only when they define the scenario's behavior or expected outcome.
5. Run `npm run e2e:fast` and iterate until green.
-
-## Migration (Angular to React Microfrontend)
-
-Pages are moving from Angular to React fragments incrementally. Today this is
narrow: the published paragraph route reads a `?react=true` flag
(`published/paragraph/paragraph.component`), the notebook footer swaps via a
`?reactFooter=true` flag (read into the notebook component's `useReactFooter`
input), the configuration table swaps via a `?reactConfiguration=true` flag
(`configuration/configuration.component`), and the notebook repository list
swaps via a `?reactNotebookRepos=true` fla [...]
-
-### Write Framework-Neutral Specs
-
-- Assert observable behavior only: what the user sees, the URL, network
effects. Avoid asserting framework internals (`[ng-version]`, Angular component
classes, `zeppelin-*` custom-element tags) except in a deliberate feature-flag
test.
-- Keep the locator order from the Locators section (role/label/text first). At
a seam that will flip frameworks, prefer a shared `data-testid` that both
implementations render.
-- Never use fixed waits at a fragment seam. Wait on a user-visible post-mount
signal or the specific remote response (`page.waitForResponse` on the fragment
chunk), then assert the rendered result. `react-footer.spec.ts` shows the
fallback pattern (`page.route('**/remoteEntry.js', route => route.abort())`).
-
-### When a Route Gains a React Flag
-
-- The flag is a route query param read via `ActivatedRoute.queryParams`, so
with the hash router it goes INSIDE the hash:
`/#/notebook/<id>/paragraph/<id>?react=true`, not before the `#`. Popups opened
by app code (`window.open`) will not carry a flag added only to `page.goto`.
-- To exercise both frameworks, follow the existing precedent and toggle the
flag in-spec: navigate the same spec with and without the flag across tests.
- - Prefer looping `for (const { label, query } of [...])` — an array of `{
label, query }` pairs — and folding `label` into the surrounding
`test.describe`/`test` name, as `notebook-repos-save-reloads-note-tree.spec.ts`
and `notebook-repo-item-workflow.spec.ts` do. The query string is the data the
test actually needs, so it travels with the label instead of being reassembled
from a bare boolean at each call site (`published-paragraph.spec.ts` predates
this and still loops a boolean; mat [...]
- - Both existing specs import the shared `NOTEBOOK_REPOS_BRANCHES` from
`e2e/models/notebook-repos-page.ts` rather than each inlining the pair — export
a same-shaped constant next to the relevant Page Object when a second spec
needs the same pair, rather than inlining it again.
- - A separate flag-appending Playwright project is an alternative, but scope
it (its own `testMatch`) to routes that read the flag rather than running the
whole suite twice.
-
-### Coverage
-
-- Coverage attribution is tracked by `PAGES` key while the denominator is
discovered from the Angular component tree. The key is the stable identity; the
path behind it is an implementation detail. While Angular still hosts the
route, keep the key mapped to that host component and cover fragment-only
behavior in the React package tests. When the Angular host component is
removed, revise the composed E2E target policy in the same change instead of
silently deleting the key. Specs keep the [...]
-
-### Suite Shape
-
-- Keep the composed suite focused on real cross-seam user flows. Behavior that
lives entirely inside one fragment belongs in that fragment's own tests; do not
grow the composed suite into a per-fragment unit suite.
-- The capture suite in `tests/notebook/core-contract/` tags its live-capture
test `@live`, because it needs an isolated Zeppelin server.
`playwright.config.js` excludes `@live`; `e2e:core-contract:live` selects it
through `playwright.core-contract.config.js` and requires an explicit server
URL. The live test deletes its own note in `finally`. The dedicated non-live
runner has no auth setup, global hooks or backend cleanup. Tag live tests
rather than excluding a whole file and hiding its [...]
-
-## Classic UI Tests (`e2e/tests/classic/`)
-
-`e2e/tests/classic/` runs Playwright against the legacy AngularJS app served
at `/classic`, ported from the retired `zeppelin-web` Protractor suite. Treat
it as a frozen legacy surface: keep it at parity coverage and test new features
only in the Angular/React suites.
-
-- **Locators (classic exception):** the classic templates predate roles and
`data-testid`, so the role/label/text-first rule cannot apply. Sanctioned here:
element ids (`#findInput`), `ng-click="..."` / `ng-controller="..."` attribute
selectors, class selectors the legacy templates already expose (`.username`,
`.interpreterHead`), and Ace/Select2 internals. Do not add `data-testid` to the
frozen `zeppelin-web` sources.
-- **Readiness:** `waitForZeppelinReady` is Angular-specific (`[ng-version]`)
and does not resolve on `/classic`; gate on a classic-visible signal instead
(e.g. the first `ParagraphCtrl` paragraph, or `.ace_text-input` attached).
-- **Coverage:** classic pages are outside the discovered Angular component
target set, so `addPageAnnotationBeforeEach` is not used here.
-- **Running:** the classic suite has its own config,
`playwright.classic.config.js` (Desktop Chrome only, targets
`http://localhost:8080`), and needs a Zeppelin server built with
`-Pweb-classic`. The `:4200` dev server does not serve `/classic`, so a plain
`npm run e2e` never includes it. Run it with `npm run e2e:classic` (single
spec: `npm run e2e:classic -- tests/classic/<spec>`). In CI the workflow
enables it on the anonymous matrix leg only
(`-Dweb.e2e.classic.disabled=false`), match [...]
-- **POM:** inlining locators/helpers is acceptable while the suite is this
small; if it grows, move them behind `models/classic-*.ts` / `*.util.ts`.
-- The React-migration / framework-neutral-spec guidance does not apply to
`tests/classic/`.