Skip to content

Added Knows Benchmark - #336

Open
farhanishmam wants to merge 2 commits into
ServiceNow:mainfrom
farhanishmam:add-knows-benchmark
Open

farhanishmam wants to merge 2 commits into
ServiceNow:mainfrom
farhanishmam:add-knows-benchmark

Conversation

@farhanishmam

Copy link
Copy Markdown
Collaborator

Added the KNOWS benchmark corresponding to the BrowserGym Knows PR #397

…s, Slides)

Knows evaluates browser agents on long-horizon document authoring in Google
Workspace. An agent drives a real browser to build a Doc, Sheet, or Slide deck
from a natural-language goal, and the result is graded checkpoint-by-checkpoint
via the Workspace APIs, yielding a fractional reward rather than binary success.

110 tasks across 22 families and 3 Workspace apps (25 docs / 45 sheets /
40 slides), exposed as a single `knows` split.

This is the AgentLab-side counterpart to the BrowserGym change that adds the
`knows` action subset, benchmark config, and task metadata. As with TimeWarp,
the benchmark itself lives in an external package (`browsergym-knows`), so
AgentLab needs only the lazy import and a docs entry:

- lazy `import browsergym.knows` in `_get_env_name`, so worker processes
  register the gym envs before `gym.make` (joblib/ray workers get a fresh
  `sys.modules`, so the parent's `prepare_backends()` import does not carry
  over)
- a `Knows` row in the supported-benchmarks table

No dependency change, matching the TimeWarp precedent: AgentLab declares no
per-benchmark extras, and benchmark packages are installed out of band per the
setup link.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
recursix
recursix previously approved these changes Jul 28, 2026

@recursix recursix left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Critical setup and reproducibility issues remain unresolved.

Review effort: Lite
Findings: 2 High severity

Open (2)
What changed in this PR

Adds AgentLab integration and documentation for the BrowserGym KNOWS benchmark.

Changes:

  • Lazily registers browsergym.knows environments.
  • Adds KNOWS to the supported benchmark table.
File Summary Findings
src/​agentlab/​experiments/​loop.py Registers KNOWS environments. Critical (4 votes): Strict reproducibility treats knows as unknown because no benchmark version case exists.
README.md Documents KNOWS benchmark setup and support. Critical (2 votes): Setup is not installable or runnable from a clean AgentLab installation. Nit (1 vote): The setup URL currently resolves to 404.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread README.md
| [MiniWoB](https://miniwob.farama.org/index.html) | [setup](https://github.com/ServiceNow/BrowserGym/blob/main/browsergym/miniwob/README.md) | 125 | Medium | 10 | no | self hosted (static files) | soon |
| [OSWorld](https://os-world.github.io/) | [setup](https://github.com/ServiceNow/AgentLab/blob/main/src/agentlab/benchmarks/osworld.md) | 369 | None | - | - | self hosted | soon |
| [TimeWarp](https://timewarp-web.github.io/) | [setup](https://github.com/sparklabutah/timewarp) | 1386 | None | 30 | yes | self hosted | soon |
| [Knows](https://alexgill321.github.io/KNOWS-benchmark/) | [setup](https://github.com/alexgill321/Agent-Benchmark) | 110 | None | 120 | yes | live web | soon |
Comment on lines +931 to +932
elif task_name.startswith("knows"):
import browsergym.knows
@farhanishmam farhanishmam self-assigned this Sep 27, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants