☰

Civitai Feed Deduplicator

Removes duplicate images/videos (same author, file name, dimensions, file size and blurhash) from Civitai feeds before they are rendered, so the masonry grid stays intact. Tested mainly on user profile pages.

Dovrai installare un'estensione come Tampermonkey, Greasemonkey o Violentmonkey per installare questo script.

You will need to install an extension such as Tampermonkey or Violentmonkey to install this script.

Dovrai installare un'estensione come Tampermonkey o Violentmonkey per installare questo script.

Dovrai installare un'estensione come Tampermonkey o Userscripts per installare questo script.

Dovrai installare un'estensione come ad esempio Tampermonkey per installare questo script.

Dovrai installare un gestore di script utente per installare questo script.

(Ho già un gestore di script utente, lasciamelo installare!)

Dovrai installare un'estensione come ad esempio Stylus per installare questo stile.

Dovrai installare un'estensione come ad esempio Stylus per installare questo stile.

Dovrai installare un'estensione come ad esempio Stylus per installare questo stile.

Dovrai installare un'estensione per la gestione degli stili utente per installare questo stile.

Dovrai installare un'estensione per la gestione degli stili utente per installare questo stile.

Dovrai installare un'estensione per la gestione degli stili utente per installare questo stile.

(Ho già un gestore di stile utente, lasciamelo installare!)

// ==UserScript==
// @name         Civitai Feed Deduplicator
// @namespace    astra.civitai.dedupe
// @version      1.3.1
// @description  Removes duplicate images/videos (same author, file name, dimensions, file size and blurhash) from Civitai feeds before they are rendered, so the masonry grid stays intact. Tested mainly on user profile pages.
// @author       Astra & V
// @license      MIT
// @match        https://civitai.com/*
// @match        https://civitai.red/*
// @run-at       document-start
// @noframes
// @grant        none
// ==/UserScript==

/*
 * How it works
 * ------------
 * Civitai loads feeds through tRPC (`image.getInfinite`). The response is
 *   { result: { data: <payload> } }
 * where <payload> is written by one of two serializers. According to the site's
 * open-source code, the server migrates superjson -> devalue pool by pool (env
 * flag), and a single response can fall back to superjson. Only devalue has
 * been observed on the live site so far:
 *   - devalue : a STRING holding a flat array. arr[0] is the root object, every
 *               value inside objects is an INDEX into the same array, -1 means
 *               undefined.
 *   - superjson: an OBJECT { json, meta }. Dates etc. are described by paths in
 *               meta ("items.3.createdAt"), so removing items requires
 *               re-numbering those paths.
 * The script wraps window.fetch, removes duplicate items from the items list
 * of either format and hands the page a normal response. Removing items before
 * the page renders them (instead of hiding them with CSS) keeps the virtualized
 * masonry grid from leaving empty gaps.
 *
 * Two items are duplicates when author id, file name, width, height, file size
 * (metadata.size) and blurhash are all identical. The first copy in feed order
 * is kept. The file size is part of the key because the blurhash only describes
 * the first frame: two different videos that start from the same frame must not
 * be mistaken for copies of each other.
 *
 * Pagination: every response carries `nextCursor`. A later request that contains
 * that cursor (in the URL or in a POST body) continues the same "chain" and
 * shares its memory of what was already shown. A request without a known cursor
 * starts a fresh chain. This does not depend on how the request input is
 * encoded, and keeps different feeds on the same page independent.
 *
 * Caveats
 * -------
 *  - Tested primarily on user profile pages. Other feeds (Most Reactions,
 *    model pages, collections) may differ and are not verified.
 *  - Only the `image.getInfinite` procedure is processed. Feeds served by other
 *    procedures are left untouched.
 *  - It relies on Civitai's internal formats (not a public API). If the site
 *    changes them, the script silently does nothing. To debug, set DEBUG = true
 *    below and look for "[dedupe]" lines in the browser console.
 *  - The superjson branch and the batch skip were tested only on synthetic data;
 *    so far only devalue responses have been observed on the live site.
 *  - Batched requests (`?batch=1`, used by the site for some logged-in
 *    cohorts) come back as a stream and are NOT processed.
 *  - Pages whose first batch of items is embedded in the HTML by the server
 *    (SSR hydration) may never go through fetch, so that first batch would not
 *    be processed. Items loaded afterwards are.
 *  - It suppresses any repeated upload from the same author, not only spam. If
 *    someone deliberately posts the same file twice, only the first one is shown.
 *
 * Privacy: the script makes no network requests of its own and sends nothing
 * anywhere. It only filters Civitai's own API responses inside the browser.
 */

(() => {
  'use strict';

  const DEBUG = false; // set to true to log removals to the console
  const log = (...a) => DEBUG && console.log('[dedupe]', ...a);

  // Only these tRPC procedures are touched.
  const PROCEDURES = ['image.getInfinite'];

  // ---------------------------------------------------------------- state --
  // One chain per feed; chains are found again through the cursors they handed out.
  const MAX_CURSORS = 500;
  const MIN_CURSOR_LENGTH = 6; // ignore tiny cursors that could match by accident
  const cursorToChain = new Map(); // nextCursor (string) -> { seen: Set<string> }

  const safeDecode = s => {
    try { return decodeURIComponent(s); } catch { return s; }
  };

  function findChain(haystack) {
    const cursors = [...cursorToChain.keys()];
    for (let i = cursors.length - 1; i >= 0; i--) { // newest first
      if (haystack.includes(cursors[i])) return cursorToChain.get(cursors[i]);
    }
    return null;
  }

  function remember(cursor, chain) {
    if (typeof cursor !== 'string' || cursor.length < MIN_CURSOR_LENGTH) return;
    cursorToChain.delete(cursor); // re-insert so it counts as newest
    cursorToChain.set(cursor, chain);
    if (cursorToChain.size > MAX_CURSORS) {
      cursorToChain.delete(cursorToChain.keys().next().value);
    }
  }

  const isPlainObject = v => v !== null && typeof v === 'object' && !Array.isArray(v);

  // ------------------------------------------------------------- dedupe core --
  function keyFrom({ type, name, hash, width, height, size, userId }) {
    if (type !== 'image' && type !== 'video') return null; // not a media item
    if (typeof name !== 'string' || typeof hash !== 'string' || !name || !hash) return null;
    // Without an author id the "same author" rule cannot be applied safely.
    if (userId === undefined || userId === null) return null;
    if (width === undefined || height === undefined) return null;
    // JSON.stringify so that a "|" inside a value cannot create false matches.
    // Missing size is allowed (older items) but then only matches another missing size.
    return JSON.stringify([userId, name, width, height, size ?? null, hash]);
  }

  // true = keep, false = duplicate. Records the key on first sight.
  function keepOnce(chain, key) {
    if (key === null) return true;
    if (chain.seen.has(key)) return false;
    chain.seen.add(key);
    return true;
  }

  // ---------------------------------------------------------- devalue format --
  const deref = (arr, i) =>
    Number.isInteger(i) && i >= 0 && i < arr.length ? arr[i] : undefined;

  function readDevalueItem(arr, idx) {
    const it = deref(arr, idx);
    if (!it || typeof it !== 'object' || Array.isArray(it)) return {};
    const user = deref(arr, it.user);
    const metadata = deref(arr, it.metadata);
    return {
      type: deref(arr, it.type),
      name: deref(arr, it.name),
      hash: deref(arr, it.hash),
      width: deref(arr, it.width),
      height: deref(arr, it.height),
      size: isPlainObject(metadata) ? deref(arr, metadata.size) : undefined,
      userId: user && typeof user === 'object' && !Array.isArray(user)
        ? deref(arr, user.id)
        : undefined,
    };
  }

  // Returns the new payload string, or null if nothing has to change.
  function processDevalue(dataStr, haystack) {
    let arr;
    try { arr = JSON.parse(dataStr); } catch { return null; }
    if (!Array.isArray(arr) || !arr[0] || typeof arr[0] !== 'object') return null;

    const root = arr[0];
    const list = deref(arr, root.items);
    if (!Array.isArray(list)) return null;

    const chain = findChain(haystack) || { seen: new Set() };
    remember(deref(arr, root.nextCursor), chain);

    const kept = list.filter(idx => keepOnce(chain, keyFrom(readDevalueItem(arr, idx))));
    if (kept.length === list.length) return null;

    log(`devalue: removed ${list.length - kept.length} of ${list.length}`);
    // The item objects stay in the array but nothing references them anymore.
    arr[root.items] = kept;
    return JSON.stringify(arr);
  }

  // -------------------------------------------------------- superjson format --
  // meta paths look like "items.3.createdAt"; re-number them after removals.
  // `keep` = old indexes of the items that survive, in order.
  // Returns the new meta, or null if its shape is not understood.
  function remapSuperjsonMeta(meta, keep) {
    const newIndex = new Map(keep.map((oldI, newI) => [oldI, newI]));
    const remapPath = p => {
      const m = /^items\.(\d+)(.*)$/s.exec(p);
      if (!m) return p;
      const n = newIndex.get(Number(m[1]));
      return n === undefined ? null : `items.${n}${m[2]}`; // null = item was removed
    };

    const out = { ...meta };

    if (meta.values !== undefined) {
      if (!isPlainObject(meta.values)) return null;
      out.values = {};
      for (const [path, annotation] of Object.entries(meta.values)) {
        const np = remapPath(path);
        if (np !== null) out.values[np] = annotation;
      }
    }

    if (meta.referentialEqualities !== undefined) {
      if (!isPlainObject(meta.referentialEqualities)) return null;
      out.referentialEqualities = {};
      for (const [first, others] of Object.entries(meta.referentialEqualities)) {
        if (!Array.isArray(others)) return null;
        // Identical objects are all written out in full, so any surviving path
        // can act as the representative.
        const paths = [first, ...others].map(remapPath).filter(p => p !== null);
        if (paths.length > 1) out.referentialEqualities[paths[0]] = paths.slice(1);
      }
    }

    return out;
  }

  // Returns the new payload object, or null if nothing has to change.
  function processSuperjson(data, haystack) {
    const json = data.json;
    if (!isPlainObject(json) || !Array.isArray(json.items)) return null;

    const chain = findChain(haystack) || { seen: new Set() };
    remember(json.nextCursor, chain);

    const keep = [];
    json.items.forEach((it, i) => {
      const fields = isPlainObject(it)
        ? {
            type: it.type,
            name: it.name,
            hash: it.hash,
            width: it.width,
            height: it.height,
            size: isPlainObject(it.metadata) ? it.metadata.size : undefined,
            userId: isPlainObject(it.user) ? it.user.id : undefined,
          }
        : {};
      if (keepOnce(chain, keyFrom(fields))) keep.push(i);
    });
    if (keep.length === json.items.length) return null;

    const out = { ...data, json: { ...json, items: keep.map(i => json.items[i]) } };
    if (data.meta !== undefined) {
      if (!isPlainObject(data.meta)) return null;
      const meta = remapSuperjsonMeta(data.meta, keep);
      if (meta === null) {
        log('superjson meta not understood, leaving the response untouched');
        return null;
      }
      out.meta = meta;
    }

    log(`superjson: removed ${json.items.length - keep.length} of ${json.items.length}`);
    return out;
  }

  function processPayload(data, haystack) {
    if (typeof data === 'string') return processDevalue(data, haystack);
    if (isPlainObject(data) && 'json' in data) return processSuperjson(data, haystack);
    return null;
  }

  // ------------------------------------------------------------ fetch hook --
  function isTargetUrl(url) {
    let path;
    try { path = new URL(url, location.href).pathname; } catch { return false; }
    return PROCEDURES.some(p => path === `/api/trpc/${p}`);
  }

  const origFetch = window.fetch;

  window.fetch = async function (...args) {
    const res = await origFetch.apply(this, args);

    try {
      const [req, init] = args;
      const url = req instanceof Request ? req.url : String(req);

      if (!isTargetUrl(url)) return res;
      if (/[?&]batch=/.test(url)) {
        log('batched request, skipping');
        return res;
      }

      // The cursor of a continuation request can be in the URL (GET) or in the
      // body (large queries are sent as POST).
      const body = init && typeof init.body === 'string' ? init.body : '';
      const haystack = `${url}\n${safeDecode(url)}\n${body}`;

      let json;
      try { json = await res.clone().json(); } catch { return res; }

      const data = isPlainObject(json) && isPlainObject(json.result) ? json.result.data : undefined;
      if (data === undefined) return res;

      const out = processPayload(data, haystack);
      if (out === null) return res;

      json.result.data = out;

      // The body is now a different size; do not keep the old size/encoding headers.
      const headers = new Headers(res.headers);
      headers.delete('content-length');
      headers.delete('content-encoding');

      return new Response(JSON.stringify(json), {
        status: res.status,
        statusText: res.statusText,
        headers,
      });
    } catch (e) {
      // Never break the page because of the userscript.
      log('error, returning the original response', e);
      return res;
    }
  };

  log('active');
})();