Skip to main content

AI Security and Safety

Text Area

This page provides a practical starting point for Genesis Mission teams--and particularly GM RFA teams--that use AI models, retrieval-augmented generation (RAG), coding assistants, or tool-using agents.

Text Area

This page does not authorize a data use or replace your institution's safety, cybersecurity, privacy, export-control, or research-security requirements. Approval to access a model or agent does not mean every project dataset may be sent to that service.

Text Area

Companion pages:

Start here: nine questions before connecting data or tools

  1. What data will the system receive? Identify the owner, sensitivity, agreements, restrictions, and approved uses.

  2. Is the exact service approved for that data? Confirm the model endpoint, tenant or workspace, retention, deletion, training use, and access terms.

  3. What can the system see? Include prompts, uploaded files, retrieved documents, embeddings, memory, logs, tool outputs, local files, and environment information.

  4. What can the system write? List read, write, delete, code execution, network access, API calls, job submission, external communication, and physical actions.

  5. Does one system combine private data, untrusted content, and a way to send data out? If the same workflow can read restricted data, takes in content you do not control (web pages, documents, email, tool output), and can communicate externally, injected instructions can make it leak that data. Remove one of the three, or put an enforced control or a person between them.

  6. What could go wrong scientifically? Consider incorrect facts, citations, code, units, parameters, assumptions, uncertainty, or experimental recommendations.

  7. Which actions require a person to approve them? Name the reviewer and define the approval point before testing.

  8. How will you stop and recover the workflow? Set time, cost, token, retry, job, and action limits. Provide cancellation and rollback.

  9. Who will you contact if something goes wrong? Identify your institution's incident reporting office and cybersecurity team, and how to reach them, before testing.

Do not proceed until the team can answer these questions.

Start at the lowest capability level that can accomplish the task. Adding tools or autonomy adds risk.

Minimum rules for every team

1. Confirm the data and service boundaries

Do not enter classified, controlled unclassified, export-controlled, proprietary, privacy-sensitive, security-sensitive, unpublished partner, or otherwise restricted information unless the data owner and the responsible institution or program have approved the exact service and workflow.

Confirm what the service retains, who can access it, whether submitted content may be used for training or service improvement, where it is processed, and how it is deleted. Use synthetic, public, or minimized data for early tests.

Outputs can be restricted too. Code, analyses, or summaries generated from controlled inputs, or that describe export-controlled or dual-use methods, may need the same handling and review as their inputs. Follow your institution's review process before publishing or sharing them.

2. Name accountable people

Record the scientific owner, technical owner, data owner, and person responsible for approving consequential actions. An AI system cannot own a decision or accept risk.

3. Use least privilege

Begin read-only. Give the system only the files, tools, data, network destinations, and operations needed for the current task. Assume a coding agent will read or write to any directory or file it can find.

Restrict network egress to the destinations the task needs. Blocking unexpected outbound connections is one of the most effective controls against data exfiltration.

4. Keep authorization outside the model

A prompt such as “never delete files” is not an access control. The application, operating system, scheduler, database, API gateway, or policy service must deny actions that are not allowed.

Require human approval before:

  • Writing to shared or authoritative data

  • Deleting or overwriting information

  • Installing software or changing an environment

  • Submitting large or expensive computing jobs

  • Sending email, posting, publishing, or contacting an external service

  • Changing permissions or credentials

  • Controlling equipment or changing an operating setpoint

  • Using an output for a safety-relevant or scientifically consequential decision

Do not approve a vague future action.

Keep approval requests rare enough that reviewers read them. If people approve a steady stream of requests without reading them, tighten permissions so fewer actions need approval.

5. Protect credentials

Never place secrets (passwords, keys, tokens, certificates, or connection strings) in prompts, notebooks, screenshots, agent instructions, committed configuration files, test fixtures, or logs.

Use an institution-approved secret service for shared or deployed systems. For local development, use approved secret management and ensure generated code or the agent sandbox cannot read the parent process environment, home directory, credential files, or cloud metadata service.

Use separate, short-lived, scope-limited credentials where supported. Know how to revoke them.

Where possible, give agents and automated workflows their own scoped identities rather than running them under a person's full account, so their actions can be limited, attributed, and revoked separately.

6. Verify scientific outputs

Treat model responses, citations, code, equations, parameters, interpretations, and conclusions as unverified.

For consequential work:

  • Compare against a non-AI baseline or trusted reference case.

  • Check citations for correctness.

  • Test generated code.

  • Check units, ranges, conservation laws, invariants, boundary conditions, and domain constraints.

  • Record uncertainty and cases where the system should abstain.

  • Require review by a qualified team member.

Agreement between two models does not establish correctness. Models can share the same errors.

7. Preserve useful evidence without leaking sensitive data

Record enough information to reproduce and investigate the workflow:

  • Model provider, endpoint, model ID, and date

  • Framework, package, and tool versions

  • Prompts or versioned instructions

  • Tool definitions and permission policy

  • Data source, version, and provenance

  • Actions, approvals, denials, errors, and timestamps

  • Resource use and cost

  • Tests, outputs, and human review decisions

Do not log secrets or unrestricted copies of sensitive prompts and data. Apply access controls, redaction, and retention rules to logs.

Assign someone to review logs and alerts, and define which events trigger an alert, such as denied actions, unexpected destinations, or limit breaches.

8. Set limits and an external kill path

Set maximum iterations, tool calls, recursion depth, retries, tokens, cost, runtime, submitted jobs, files changed, and external actions. Stop when a limit is reached rather than asking the model whether it should continue.

An external kill path stops the workflow without relying on the model: for example, killing the process, revoking its credential, or disabling its service account.

9. Vet third-party tools, servers, skills, and models

Treat MCP servers, agent skills, plugins, IDE extensions, and downloaded model weights as software that runs with your permissions. Before connecting one:

  • Get it from a source your institution trusts, and record its version or commit.

  • Review what it can read, write, execute, and reach on the network.

  • Pin the version, and review it again before each update.

  • Run local servers in a sandbox or container with only the access the task needs.

  • For model weights, verify checksums or signatures, and prefer formats that cannot execute code when loaded, such as safetensors.

  • Follow your institution's policy on which open-weight models and model sources are allowed.

Keep a list of approved tools, servers, skills, and models for each project.

10. Treat retrieved content, memory, and other agents' messages as untrusted

Anything the system reads can carry instructions: documents, web pages, tool output, its own stored memory, and messages from other agents. Handle all of it as data, not as commands.

  • Limit who and what can write to memory and retrieval corpora, and record the source of each entry.

  • Review or expire stored memory so that content from one session cannot quietly steer later sessions.

  • In multi-agent systems, authenticate agents to each other and give each agent only the permissions its own task needs. A request from another agent does not carry that agent's authority.

  • Content from any of these sources must not trigger a high-impact action without the approval required in rule 4.

11. Keep physical safety independent of the AI system

When an AI system can recommend or change settings on instruments, experiments, or facility equipment:

  • Keep safety interlocks, limits, and emergency stops in hardware or in control-system logic that the AI system cannot change.

  • Define the safe operating envelope before testing, and enforce it outside the model.

  • Run new workflows in simulation or dry-run mode before connecting them to equipment.

  • Involve the facility's safety and operations staff before the AI system is connected.

  • Make sure an operator can take manual control at any time.

12. Test before shared or operational use

At minimum, test that the system:

  • Refuses or blocks access to another project or user's data

  • Cannot call an unauthorized tool

  • Cannot turn a read operation into a write

  • Requires valid approval for high-impact actions

  • Cannot be made to take an unauthorized action, access unauthorized data, or send data externally by instructions embedded in retrieved content, tool output, stored memory, or messages from other agents

  • Stops at resource, cost, time, and retry limits

  • Does not expose secrets in output or logs

  • Produces acceptable results on representative scientific tasks

  • Fails safely when a model, tool, network service, or data source is unavailable

Repeat tests after material changes to the model, prompts, tools, retrieval corpus, memory, permissions, or framework.

Text Area

Pre-deployment checklist

Before a workflow is shared, deployed, or connected to consequential resources, retain:

















Record these in the agent card template and model card template on the AI and Agents page!

Text Area

What to do when something goes wrong

Examples include:

  • A secret appears in a prompt, output, repository, screenshot, or log

  • The system accesses another user or project's data

  • An agent takes or attempts an unapproved action

  • Retrieved content or tool output changes the agent's goal

  • A job, loop, API call, or cost runs beyond its limit

  • A model or tool sends data to an unexpected destination

  • An incorrect AI result influences a consequential decision

  • Equipment behaves unexpectedly after an AI recommendation or action

Response steps:

  1. Stop or disable the workflow. Use the external kill path; do not rely on the model or agent to stop itself.

  2. Prevent additional actions. Isolate the service, tool, credential, job, or equipment connection as appropriate.

  3. Revoke exposed credentials. Rotate keys and tokens using the approved process.

  4. Preserve evidence. Keep relevant prompts, logs, tool calls, files, model and framework versions, approvals, and timestamps. Do not destroy evidence while cleaning up.

  5. Report the event. Notify your institution's incident reporting office and cybersecurity team, and its safety office if equipment or people were affected. Follow your institution's required reporting process and timelines.

  6. Do not resume until the cause, impact, corrective action, and required approvals are reviewed.

Do not include secrets or unnecessarily sensitive data in a help ticket or chat message.

Security and evaluation tools

Tools support review; they do not certify a system or replace data approval, access control, scientific testing, or human judgment.

Tool or practiceUse it forDo not treat it as
Trivy, Snyk, Dependabot, or institution-approved equivalentsFinding known vulnerable dependencies, containers, and packagesA test of prompt injection, agent permissions, data handling, or scientific correctness
Repository secret scanningFinding credentials in source control; enable pre-commit hooks and push protection to block them before they are committedProof that secrets cannot leak through runtime access, prompts, memory, or logs
Evaluation and red-team harnesses (for example, Promptfoo, PyRIT, garak)Repeatable evaluation and adversarial testing of model, RAG, agent, or API behaviorA source-code scanner, authorization system, or guarantee that all attacks are covered
Unit, integration, and scientific testsChecking expected behavior, invariants, reference cases, and regressionsA replacement for adversarial testing or operational monitoring

Evaluation and red-team harnesses

Open-source harnesses such as Promptfoo, Microsoft's PyRIT, and NVIDIA's garak run repeatable evaluations and adversarial tests against model, RAG, agent, and API targets. They take different approaches: Promptfoo runs configuration-driven test suites, PyRIT supports scripted multi-turn attack campaigns, and garak scans with a library of probes. No single harness covers every attack.

Use only approved test data and endpoints. Review generated tests before execution, especially when the target can call tools or external services.

Check where a harness sends data before you run it. Some collect usage telemetry or generate attack content through a vendor-hosted service by default; Promptfoo does both. Turn these features off unless that service is approved for your project's data.

Promptfoo was acquired by OpenAI in 2026 and remains open source. Teams evaluating OpenAI models should consider adding an independent harness.

Documentation:

What not to rely on

Do not rely on any of the following as the only protection:

  • A system prompt that tells the model to behave

  • XML tags, delimiters, or keyword filters

  • A model's confidence statement

  • Agreement between multiple models

  • A red-team scan with no human review

  • Logs that no one monitors

  • A human approval button that does not show the exact action and parameters

Reference resources

Frameworks and standards

NIST is revising the AI RMF under the White House AI Action Plan; check for a newer version before relying on the 2023 text. Also in development and worth watching: the Cyber AI Profile (NIST IR 8596, preliminary draft December 2025), SP 800-53 control overlays for AI systems including single-agent and multi-agent deployments (COSAiS), and agent-focused guidance from NIST's Center for AI Standards and Innovation (CAISI).

Threat catalogs and practitioner guidance

Government security guidance

Text Area

Submit a support ticket for any of the following:

  • Connect your team with services like data management planning, supercharging your scientific workflows with AI best practices, and cross-cutting AI capabilities.

  • Have an issue, suggestion, or addition to this content.

Click to Submit a Support Ticket