Elite Personas LLC operates AI-generated personas — fictional characters generated by our own software. This page describes the safety controls built into our systems: how the imagery we publish is screened, how our personas are disclosed as AI, and how our direct-message system is designed to converse safely. We publish it because an AI-content operation should be able to state plainly what its safeguards are, and be held to them.
Status of our direct-message (chatbot) system. Our personas' ability to hold direct-message conversations is built but currently switched off, pending legal review. In its current state it cannot send a message to anyone. Our intended operating model is autonomous: once enabled, a persona composes and sends its own replies without a human approving each message — so the automated safeguards described below, together with after-the-fact human review of conversation logs, are the controls we rely on. We describe them here honestly, including their limits.
For ordinary conversation, a persona's reply is generated by a language model constrained by fixed conduct rules and screened by automated filters before it is sent. For a defined set of safety-sensitive situations, the system does not let the language model answer freely at all — it uses fixed, pre-written responses instead (Section 3). Because the intended model is autonomous, these automated controls — not a human reading each message — are what stand between a user and a reply.
Incoming messages are screened by automated classifiers before a persona composes a reply — including checks for manipulation and prompt-injection attempts, toxicity, and attempts to extract sensitive data, and a dedicated pattern set that looks for signals a sender may be a minor (Section 4).
A generated reply is independently screened before sending — checking for toxicity, leaked personal information, and content-policy violations — and higher-risk cases are escalated to a specialized safety classifier that flags categories such as violence, non-consensual or exploitative sexual content, hate, and self-harm; an unreadable result from that classifier is treated as unsafe. If a screening service is unavailable, the reply is held rather than sent unscreened. These automated filters reduce risk but are not represented as perfect (see Section 8).
For a defined set of sensitive incoming messages, the system bypasses the language model entirely and responds with a fixed, pre-written response rather than generated text. Those categories are: crisis / self-harm, a direct "are you AI?" question, requests to meet in person or move off-platform, requests for medical, legal, or financial advice, scam / money solicitation, and any signal that a user may be a minor. The detection and routing for these categories is built and active. The exact wording of these responses is under review by counsel: the placeholder responses are marked as unreviewed and are not released for sending, and the messaging feature as a whole is currently switched off.
A dedicated pattern set scans every incoming message — before any language model sees it — for signals that a sender may be under 18 (age claims, concealment language, school-grade context, and similar). It is intentionally biased toward over-flagging. On a hit, the system is designed to respond only with the minor-handling response, end the conversation, and set a block on that user that persists for the rest of the conversation and does not reset.
This is content-based detection of minor signals within a conversation. It is not identity or age verification of a subscriber — on the paid platform, age and identity verification of account holders is handled by that platform as a condition of its service. We do not represent that we independently verify anyone's age.
Every image a persona publishes passes automated safety gates before it can be posted.
An automated apparent-age classifier runs on every image of a persona, at every content tier, and is the authority on age. It fails closed: if the check cannot run, the image is rejected. No such image is published without passing it.
Before a scene is rendered, an automated gatekeeper reviews it for underage implications, non-consensual framing, trademark or copyright issues, and platform-policy violations; any ambiguity about age is resolved as a rejection. Generated media is then classified and held to its declared content tier, with structural and identity checks layered on. Content that fails is quarantined, not published.
The reference material our systems draw on is filtered by deterministic safety lists that remove exploitative and prohibited categories. Those filters fail closed — if a list cannot load, it blocks rather than allows.
Every persona's public social-media profile carries a fixed AI-disclosure line, applied automatically by the software that writes the profile. There is no opt-out — the software re-applies the disclosure on every profile write, and if a bio would exceed the platform's length limit the persona's own text is trimmed, never the disclosure. Every one of our live personas carries it today.
AI-generated images and video published by our channels are labeled as AI-generated, and adult-oriented media is self-labeled using each platform's official content-labeling mechanism.
If a user asks a persona whether it is AI, the system routes to a fixed AI-acknowledgment response rather than denying it. Our personas are fictional characters and are not represented as real people.
Because the intended model is autonomous, oversight is continuous review of what the system actually said: direct-message conversations are logged and reviewed internally for safety, rather than approved message-by-message before sending.
Every automated component has an operator-controlled off-switch — overall, and per platform for the components that post or message — so any activity can be halted immediately. Monetization and conversion features are off by default and require a deliberate action to enable.
Safety-relevant events — screening blocks, sensitive-topic routing, and message sends — are recorded to timestamped logs so our conduct can be reviewed after the fact.
Conversation and engagement logs are retained for 30 days for internal safety review, then automatically deleted. We do not keep this data indefinitely.
No automated safety system is perfect. Our safety gates are designed to fail toward caution — to block when in doubt, and to hold a message rather than send it if a screening service is unavailable — and our conduct is reviewable through the logs we keep. But we do not represent these systems as infallible: automated classifiers can be wrong, and language-model instructions can in principle be circumvented, which is why we layer independent checks rather than rely on any single control. We do not independently verify the age of platform users. We welcome reports of any concern about our channels' conduct at the address below, and we act on them.
To report a safety concern, or for questions about the safeguards described here:
Elite Personas LLC
Email: admin@elitepersonas.com