docs(18): capture phase context
This commit is contained in:
parent
4decefb940
commit
a37dcebd2d
2 changed files with 314 additions and 0 deletions
|
|
@ -0,0 +1,229 @@
|
||||||
|
# Phase 18: Campaign Grouping & Phishing Analysis API - Context
|
||||||
|
|
||||||
|
**Gathered:** 2026-07-15
|
||||||
|
**Status:** Ready for planning
|
||||||
|
|
||||||
|
<domain>
|
||||||
|
## Phase Boundary
|
||||||
|
|
||||||
|
Duplicate reports of the same phishing/spam campaign are automatically
|
||||||
|
grouped into a shared `campaigns` row (tiered key: Message-ID, then
|
||||||
|
attachment-hash/URL-domain + subject + sender + time-window, then sender +
|
||||||
|
normalized subject + client + time-window), accumulating linked reports and
|
||||||
|
recipients over time. An operator can trigger on-demand analysis of one
|
||||||
|
specific ticket (`POST /api/phishing/tickets/{ticket_id}/analyze`) or browse
|
||||||
|
campaigns (`GET /api/phishing/campaigns`, `GET /api/phishing/campaigns/{id}`)
|
||||||
|
through a properly access-controlled `/api/phishing/*` surface — the first
|
||||||
|
API routes in this milestone, establishing the auth convention every later
|
||||||
|
phishing endpoint (Phases 19-21) must also follow.
|
||||||
|
|
||||||
|
</domain>
|
||||||
|
|
||||||
|
<decisions>
|
||||||
|
## Implementation Decisions
|
||||||
|
|
||||||
|
### Grouping Trigger Point
|
||||||
|
- **D-01:** Campaign grouping runs automatically as part of the same
|
||||||
|
detection path Phase 15's `phishing-detector.ts` already uses (webhook
|
||||||
|
fire-and-forget `ticket.created` handler + the scheduled cron sweep) — the
|
||||||
|
moment a `reports` row is created or updated, grouping runs against it so a
|
||||||
|
campaign forms/accumulates without waiting for anyone to call the API. The
|
||||||
|
new `POST /api/phishing/tickets/{ticket_id}/analyze` endpoint (DETECT-03)
|
||||||
|
calls the SAME shared grouping function on-demand for one ticket — not a
|
||||||
|
separate/duplicate implementation. Rationale: success criterion #2
|
||||||
|
("accumulates... as new duplicate reports arrive over time") implies
|
||||||
|
grouping must happen automatically, not only when explicitly requested.
|
||||||
|
|
||||||
|
### Grouping Parameters
|
||||||
|
- **D-02:** Time window for the fallback tiers (attachment-hash/URL-domain +
|
||||||
|
subject + sender, and sender + normalized-subject + client) is **24
|
||||||
|
hours**. Rationale: catches same-day mass-phishing blasts (the realistic
|
||||||
|
case — one attacker campaign, multiple people reporting the same day)
|
||||||
|
without over-grouping unrelated reports that happen to share sender+subject
|
||||||
|
weeks apart.
|
||||||
|
- **D-03:** Subject normalization for the fallback tiers: strip leading
|
||||||
|
`Re:`/`Fwd:`/`Fw:` prefixes (case-insensitive, repeated occurrences),
|
||||||
|
lowercase, trim whitespace. Standard email-threading normalization, nothing
|
||||||
|
more elaborate.
|
||||||
|
|
||||||
|
### Campaign Merge Behavior
|
||||||
|
- **D-04:** Campaigns are never merged into each other. Each new report
|
||||||
|
attaches to at most one existing campaign (the single best/first match by
|
||||||
|
the tiered key) or creates a new one. If a report's tiered key
|
||||||
|
theoretically matches two distinct existing open campaigns, this is
|
||||||
|
treated as a rare edge case — do NOT implement multi-campaign-merge logic
|
||||||
|
(transactional reassignment of reports/messages/indicators/classifications
|
||||||
|
across campaigns) in this phase. Attach to the first/best match found and
|
||||||
|
move on; revisit only if this proves to be a real, recurring problem in
|
||||||
|
practice.
|
||||||
|
|
||||||
|
### Permission Model
|
||||||
|
- **D-05:** Add a new `phishing` resource to `lib/permissions.ts`'s
|
||||||
|
`statement` with the FULL action set from the milestone spec now —
|
||||||
|
`phishing: ["read", "analyze", "approve", "remediate"]` — but this phase
|
||||||
|
only GRANTS `read` and `analyze` in role definitions:
|
||||||
|
- `superAdminRole` / `adminRole`: `phishing: ["read", "analyze"]`
|
||||||
|
- `userRole`: `phishing: ["read"]` (can browse campaigns, cannot trigger
|
||||||
|
on-demand analysis)
|
||||||
|
`approve`/`remediate` are declared in the statement (so the vocabulary
|
||||||
|
exists) but ungranted to any role until Phase 20 wires them up. Rationale:
|
||||||
|
establishes the full permission vocabulary once so Phase 20 only adds role
|
||||||
|
grants, not new statement keys — avoids touching the same statement block
|
||||||
|
twice across phases.
|
||||||
|
|
||||||
|
### Route Auth Granularity
|
||||||
|
- **D-06:** Every `/api/phishing/*` route uses the fine-grained
|
||||||
|
`requirePermission()` check from day one, not plain `requireAuth()`:
|
||||||
|
- `GET /api/phishing/campaigns` and `GET /api/phishing/campaigns/{id}` →
|
||||||
|
`requirePermission('phishing', 'read')`
|
||||||
|
- `POST /api/phishing/tickets/{ticket_id}/analyze` →
|
||||||
|
`requirePermission('phishing', 'analyze')`
|
||||||
|
This establishes the exact auth convention (per-route, per-action
|
||||||
|
permission check) that Phase 19-21's endpoints must copy. Since `userRole`
|
||||||
|
gets `read` granted per D-05, this doesn't restrict any current user from
|
||||||
|
browsing — it wires the fine-grained check through now rather than
|
||||||
|
retrofitting it later.
|
||||||
|
|
||||||
|
### Claude's Discretion (explicitly deferred to research + planner)
|
||||||
|
- **Exact SQL/query shape for `GET /api/phishing/campaigns/{id}`'s nested
|
||||||
|
response** (linked reports, messages, indicators, classification history)
|
||||||
|
— single JOIN-heavy query vs. multiple queries assembled in application
|
||||||
|
code. Follow whatever the existing `entity-sync.ts` / DetailModal-backing
|
||||||
|
API routes in this codebase already do for similarly-shaped nested detail
|
||||||
|
responses.
|
||||||
|
- **Exact shape of the tiered-key matching queries** (how "attachment-hash/
|
||||||
|
URL-domain" is queried against the `indicators` table's `indicator_type`
|
||||||
|
values `'attachment_hash'`/`'url'`/`'sender'` from Phase 16, how URL-domain
|
||||||
|
is extracted from a full URL string) — implementation detail, resolve via
|
||||||
|
research against the actual `indicators` schema and data shapes Phase 16
|
||||||
|
produces.
|
||||||
|
- **Response body shape for `POST /analyze`** — should return the resulting
|
||||||
|
campaign linkage (campaign ID, grouping method used, whether a new
|
||||||
|
campaign was created vs. an existing one was matched) — exact field names
|
||||||
|
are planner's call, following the project's camelCase API response
|
||||||
|
convention.
|
||||||
|
- **Where the shared grouping function lives** (new file e.g.
|
||||||
|
`lib/services/campaign-grouping-service.ts` vs. extending
|
||||||
|
`phishing-detector.ts`) — planner's call, following whatever composition
|
||||||
|
pattern is cleanest given `phishing-detector.ts`'s current size and
|
||||||
|
responsibilities.
|
||||||
|
- **Whether `campaigns.status` transitions in this phase** (e.g., does
|
||||||
|
grouping ever set/change `status` beyond the migration's default `'open'`)
|
||||||
|
— not mentioned in success criteria; likely out of scope for this phase
|
||||||
|
(status transitions are Phase 19/20's classification/remediation concern),
|
||||||
|
but confirm during planning.
|
||||||
|
|
||||||
|
</decisions>
|
||||||
|
|
||||||
|
<canonical_refs>
|
||||||
|
## Canonical References
|
||||||
|
|
||||||
|
**Downstream agents MUST read these before planning or implementing.**
|
||||||
|
|
||||||
|
### Phase 15 detection path (grouping hooks into this)
|
||||||
|
- `lib/services/phishing-detector.ts` — `detectPhishingTicket()`,
|
||||||
|
`gatherTicketEvidence()`, the `content_hash` idempotency pattern, and the
|
||||||
|
`reports` upsert (`ON CONFLICT (ticket_id) DO UPDATE`) this phase's
|
||||||
|
grouping call attaches to.
|
||||||
|
- `lib/services/webhook-service.ts` — the `ticket.created` fire-and-forget
|
||||||
|
hook where Phase 15 wired in detection; this phase's grouping call should
|
||||||
|
follow the same fire-and-forget wiring point (per D-01).
|
||||||
|
- `lib/services/sync-scheduler.ts` — the cron sweep pattern (`pax8-daily`-
|
||||||
|
style schedule row) Phase 15 added for reconciliation; grouping runs on
|
||||||
|
the same sweep per D-01.
|
||||||
|
|
||||||
|
### Phase 16 evidence (grouping keys off this)
|
||||||
|
- `lib/services/phishing-eml-service.ts` — writes `indicators` rows with
|
||||||
|
`indicator_type` values `'attachment_hash'`, `'url'`, `'sender'` (confirmed
|
||||||
|
by reading the file directly) — the CAMP-01 fallback-tier keys come from
|
||||||
|
these rows.
|
||||||
|
- `lib/services/eml-parser.ts` — `NormalizedMessage`'s `Message-ID` field
|
||||||
|
(via `parseEml`) — the CAMP-01 primary-tier key.
|
||||||
|
|
||||||
|
### Auth/permissions (this phase establishes the phishing resource here)
|
||||||
|
- `lib/permissions.ts` — `statement`, `superAdminRole`, `adminRole`,
|
||||||
|
`userRole` — add the new `phishing` resource here (D-05). Existing
|
||||||
|
`rmm: ["read", "execute"]` resource is the closest analog for a
|
||||||
|
read/action-verb permission shape.
|
||||||
|
- `lib/auth-utils.ts` — `requirePermission(resource, action)` signature this
|
||||||
|
phase's routes call (D-06). See `app/api/admin/rmm/settings/route.ts` for
|
||||||
|
the exact call-and-early-return pattern:
|
||||||
|
`const { error } = await requirePermission('rmm', 'read'); if (error) return error;`
|
||||||
|
|
||||||
|
### Schema (read/write for this phase)
|
||||||
|
- `migrations/097_phishing_triage_schema.sql` — `campaigns` table
|
||||||
|
(`campaign_key`, `group_method`, `first_seen_at`, `last_seen_at`,
|
||||||
|
`report_count`, `status`), `reports.campaign_id` (nullable FK, populated
|
||||||
|
starting this phase per the migration's own comment).
|
||||||
|
|
||||||
|
### Prior phase decisions (for consistency)
|
||||||
|
- `.planning/phases/15-data-model-detection-ticket-evidence/15-CONTEXT.md`
|
||||||
|
(D-01: webhook + cron "primary path, cron reconciles" pattern; D-04:
|
||||||
|
content-hash idempotency) — this phase's grouping trigger follows the same
|
||||||
|
dual-path wiring instinct.
|
||||||
|
- `.planning/phases/17-mimecast-blast-radius-lookup/17-CONTEXT.md` (D-05:
|
||||||
|
single-tenant-for-v1, documented-limitation-over-silent-gap pattern) —
|
||||||
|
shows the project's general preference for "simple now, documented
|
||||||
|
limitation, defer complexity" when a rare edge case (like campaign
|
||||||
|
merging) would add substantial implementation cost.
|
||||||
|
|
||||||
|
</canonical_refs>
|
||||||
|
|
||||||
|
<code_context>
|
||||||
|
## Existing Code Insights
|
||||||
|
|
||||||
|
### Reusable Assets
|
||||||
|
- `phishing-detector.ts`'s `computePhishingContentHash()` — precedent for a
|
||||||
|
small pure hashing/normalization helper; the new subject-normalization
|
||||||
|
function (D-03) should follow the same style.
|
||||||
|
- `requirePermission()` from `lib/auth-utils.ts` — used as-is, no changes
|
||||||
|
needed to the helper itself, only new resource/action strings.
|
||||||
|
|
||||||
|
### Established Patterns
|
||||||
|
- Factory/service convention: pure functions in `lib/services/`, called from
|
||||||
|
both a webhook handler and a cron-scheduled sync route — same shape
|
||||||
|
`phishing-detector.ts` already established, this phase extends it rather
|
||||||
|
than inventing a new wiring pattern.
|
||||||
|
- `ON CONFLICT ... DO UPDATE` upsert pattern (`reports` table) — the
|
||||||
|
campaign accumulation logic (D-01, CAMP-02) likely needs a similar
|
||||||
|
upsert-or-create-then-link pattern against `campaigns`.
|
||||||
|
|
||||||
|
### Integration Points
|
||||||
|
- No `/api/phishing/*` routes exist yet in this repo — this phase creates
|
||||||
|
the first ones. No existing route to pattern-match against inside this
|
||||||
|
specific path; use `app/api/admin/rmm/settings/route.ts` (or similar
|
||||||
|
`requirePermission`-gated routes) as the general Next.js route + auth
|
||||||
|
pattern instead.
|
||||||
|
|
||||||
|
</code_context>
|
||||||
|
|
||||||
|
<specifics>
|
||||||
|
## Specific Ideas
|
||||||
|
|
||||||
|
No specific UI/output-format requirements — this is a backend abstraction +
|
||||||
|
API surface with no UI in this phase (`UI hint: no` per ROADMAP.md).
|
||||||
|
|
||||||
|
</specifics>
|
||||||
|
|
||||||
|
<deferred>
|
||||||
|
## Deferred Ideas
|
||||||
|
|
||||||
|
- **Multi-campaign merge logic** — deferred per D-04. If a report's tiered
|
||||||
|
key ever legitimately matches two distinct existing campaigns in practice
|
||||||
|
(not just in theory), building proper merge/reassignment logic is future
|
||||||
|
work, not blocking this phase.
|
||||||
|
- **`campaigns.status` transition logic** (open → closed/resolved etc.) —
|
||||||
|
out of scope for this phase; belongs to Phase 19 (classification) or
|
||||||
|
Phase 20 (remediation).
|
||||||
|
- **approve/remediate permission grants** — the `phishing` resource's
|
||||||
|
`approve`/`remediate` actions are declared in the statement (D-05) but not
|
||||||
|
granted to any role until Phase 20.
|
||||||
|
|
||||||
|
None — discussion otherwise stayed within phase scope.
|
||||||
|
|
||||||
|
</deferred>
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
*Phase: 18-campaign-grouping-phishing-analysis-api*
|
||||||
|
*Context gathered: 2026-07-15*
|
||||||
|
|
@ -0,0 +1,85 @@
|
||||||
|
# Phase 18: Campaign Grouping & Phishing Analysis API - Discussion Log
|
||||||
|
|
||||||
|
> **Audit trail only.** Do not use as input to planning, research, or execution agents.
|
||||||
|
> Decisions are captured in CONTEXT.md — this log preserves the alternatives considered.
|
||||||
|
|
||||||
|
**Date:** 2026-07-15
|
||||||
|
**Phase:** 18-campaign-grouping-phishing-analysis-api
|
||||||
|
**Areas discussed:** Grouping trigger, Grouping params, Campaign merge, Permission model, Route auth
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Grouping Trigger
|
||||||
|
|
||||||
|
| Option | Description | Selected |
|
||||||
|
|--------|-------------|----------|
|
||||||
|
| Auto-group on every detection | Hook grouping into the same webhook + cron path Phase 15 uses; new `/analyze` endpoint reuses the same shared function on-demand. | ✓ |
|
||||||
|
| On-demand only via the new API | Grouping only happens when `POST /analyze` is explicitly called. | |
|
||||||
|
|
||||||
|
**User's choice:** Auto-group on every detection (recommended option).
|
||||||
|
**Notes:** None.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Grouping Params
|
||||||
|
|
||||||
|
| Option | Description | Selected |
|
||||||
|
|--------|-------------|----------|
|
||||||
|
| 24h window, strip Re:/Fwd: + lowercase + trim | Same-day mass-phishing blast window; standard subject normalization. | ✓ |
|
||||||
|
| 72h window, same subject normalization | Wider window, catches slower-trickling reports at cost of looser grouping. | |
|
||||||
|
| You decide (Claude's discretion) | No strong preference. | |
|
||||||
|
|
||||||
|
**User's choice:** 24h window, strip Re:/Fwd: + lowercase + trim (recommended option).
|
||||||
|
**Notes:** None.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Campaign Merge
|
||||||
|
|
||||||
|
| Option | Description | Selected |
|
||||||
|
|--------|-------------|----------|
|
||||||
|
| Never merge existing campaigns | Attach to single best-matching existing campaign or create new; no merge logic. | ✓ |
|
||||||
|
| Support merging two campaigns when both match | Transactional reassignment of reports/messages/indicators/classifications across campaigns. | |
|
||||||
|
|
||||||
|
**User's choice:** Never merge existing campaigns (recommended option).
|
||||||
|
**Notes:** None.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Permission Model
|
||||||
|
|
||||||
|
| Option | Description | Selected |
|
||||||
|
|--------|-------------|----------|
|
||||||
|
| New 'phishing' resource: read + analyze actions now, full statement vocabulary declared | `phishing: ["read","analyze","approve","remediate"]` in statement; only read+analyze granted this phase. | ✓ |
|
||||||
|
| New 'phishing' resource: read + analyze actions only, extend later | Only declare read+analyze now; Phase 20 adds approve/remediate to the statement itself. | |
|
||||||
|
|
||||||
|
**User's choice:** New 'phishing' resource: read + analyze actions now (recommended option).
|
||||||
|
**Notes:** None.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Route Auth (follow-up)
|
||||||
|
|
||||||
|
| Option | Description | Selected |
|
||||||
|
|--------|-------------|----------|
|
||||||
|
| requirePermission('phishing','read'/'analyze') on every route | Fine-grained check on all routes from day one. | ✓ |
|
||||||
|
| requireAuth() only for GETs, requirePermission for POST /analyze | Looser check on reads. | |
|
||||||
|
|
||||||
|
**User's choice:** requirePermission on every route (recommended option).
|
||||||
|
**Notes:** Asked as a follow-up after the Permission Model question to nail down route-level specifics.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Claude's Discretion
|
||||||
|
|
||||||
|
- Exact SQL/query shape for `GET /api/phishing/campaigns/{id}`'s nested response
|
||||||
|
- Exact shape of the tiered-key matching queries against the `indicators` table
|
||||||
|
- Response body shape for `POST /analyze`
|
||||||
|
- Where the shared grouping function lives (new file vs. extending `phishing-detector.ts`)
|
||||||
|
- Whether `campaigns.status` transitions in this phase (likely out of scope)
|
||||||
|
|
||||||
|
## Deferred Ideas
|
||||||
|
|
||||||
|
- Multi-campaign merge logic — deferred per D-04, future work if it proves to be a real recurring problem
|
||||||
|
- `campaigns.status` transition logic — belongs to Phase 19/20
|
||||||
|
- approve/remediate permission grants — declared now (D-05), granted in Phase 20
|
||||||
Loading…
Add table
Add a link
Reference in a new issue