Documentation

Build your first survey in minutes

Everything you need to prepare, publish, and share a SuAVE survey — from CSV formatting to Jupyter notebook integration.

Start here

What is SuAVE?

SuAVE (Survey Analysis via Visual Exploration) is a free web platform that turns a CSV file into an interactive, multi-view data explorer for images, maps, charts, and networks. Upload your data, add qualifiers where needed, and SuAVE renders a zoomable grid of image tiles linked to a faceted filter panel — no installation, no code, no database configuration. The platform is hosted at suave.sdsc.edu and is free for academic and non-commercial use.

The open-source SuAVE 2026 frontend (React 18 + Vite) can also be self-hosted and loaded with any CSV directly from the browser or via URL parameters — no account required for local exploration.

The five-minute workflow

  1. Prepare a CSV where each row is one survey item (respondent, specimen, artwork, location, …)
  2. Annotate column names with #number, #multi, #date, etc. to control how they appear
  3. If you have images, add them to the survey via the New Survey dialog (see Images section)
  4. Log in at suave.sdsc.edu, click New Survey, and upload the CSV
  5. Share the resulting URL — anyone with it can explore your data without an account
💡

SuAVE works best with 50–50,000 rows. Very small datasets (under 20 items) have limited visual value; very large ones (100,000+) may require server-side pagination — contact us for large-data deployments.

Additional resources

The full reference documentation is at suave-ucsd.github.io/SuAVE-Documentation, covering account setup, image galleries, LimeSurvey integration, bibliographic networks, Jupyter notebooks, and self-hosting. Video tutorials are on the SuAVE YouTube channel.

CSV format

SuAVE reads standard comma-separated values. The first row is the header — it defines ordinary variable names, reserved variable names, and optional qualifier suffixes (see next section). Values can be text, numbers, dates, pipe-separated lists, or URLs.

Minimal example

# Every row is one survey item. Columns can be in any order. #name, Species#multi, Period#number, #img, Notes#long, Latitude#number#hidden, Longitude#number#hidden "Ammonite A", "Cephalopoda|Mollusca", 145, "amm001", "Well-preserved suture lines", 51.5, -0.12 "Belemnite B", "Cephalopoda", 160, "bel002", "", 52.1, 0.34

Rules

For an ordinary variable, the display label is the part before the first #; any following tokens are qualifiers. A column named Geological Period#number#hidden has the label "Geological Period", is treated as a numeric range slider, and is hidden from the filter panel (but shown in the item detail panel). Reserved variable names such as #name are the exception: the entire header is the variable name.

Values in #multi columns are pipe-separated: Cephalopoda|Mollusca|Marine. Each pipe-delimited token becomes an independent checkbox in the filter panel.

UTF-8 encoding is required. Quote fields that contain commas. Empty cells are treated as missing values and displayed as "N/A" in the filter panel. Grouped views (Bucket, Bars, Crosstab and the charts) and the map leave items with no value out by default; tick Display missing values in Settings to show them as an N/A group. On the map they are always drawn, in neutral grey.

Loading options

Beyond direct upload, SuAVE 2026 accepts data from several sources via URL parameters. A Google Sheets share link (the /edit or /view form) is automatically converted to its CSV export equivalent. Any publicly accessible CSV can be loaded with ?csv_url=URL. A survey config JSON file can bundle the CSV URL, DZC URL, and view settings into a single shareable document (see Survey config files).

Variable naming and qualifier conventions

SuAVE CSV headers use three distinct conventions: exact reserved variable names, reserved words in spatial variable labels, and qualifiers appended to ordinary variable names. These conventions are not interchangeable.

Reserved variable names

The following are complete, exact column names. They begin with #, but they are variables—not qualifiers to append to another label.

Variable name Purpose Value in each row
#nameItem display name and detail-panel titlePlain text
#imgPrimary image identifierImage ID matching the survey's DZI/DZC data
#hrefMakes the #name title clickableItem-level URL
#netvisAssociates the item with network-visualization dataNetwork-data reference used by the Netvis view

#uniqueid and #views are legacy/internal headers recognized by the original SuAVE loader. Do not add them to new SuAVE 2026 CSV files: row IDs are generated automatically, and enabled views are configured in the survey settings.

ℹ️

Use these exact headers in new CSV files: #name, #img, and #href—not Name#name, Image#img, or Website#href. SuAVE 2026 accepts some labeled forms for compatibility, but the canonical SuAVE convention and the original implementation use exact reserved variable names.

Map columns

A column tagged #lat, #lon or #geometry is the map's latitude, longitude or shape column, whatever its label says (Широта#lat, Route#geometry#hiddenmore). Without these tags the map looks, as the original SuAVE did, for a label (the part before the first #) containing one of the words below, case-insensitively, so existing surveys keep working. #coord places nothing.

Or a label containing Purpose Recommended header
LatitudePoint latitudeLatitude#number#hidden
LongitudePoint longitudeLongitude#number#hidden
GeometryWKT or GeoJSON geometry for points, lines, or polygonsGeometry#hiddenmore

Column qualifiers

Qualifiers are suffixes appended to an ordinary variable name with #. They control the data type, filtering, sorting, and visibility. Multiple compatible qualifiers can be chained, as in Depth#number#hidden.

Qualifier Filter panel Detail panel Notes
(none)Checkbox list✓Default for text columns
#numberRange slider + histogram✓Values must be numeric
#dateDate range picker✓ISO 8601 or MM/DD/YYYY
#multiCheckbox list✓Pipe-separated values; each token is a separate facet
#longSearch scope only✓Long text — appears in search scope dropdown, not as checkboxes
#link—Clickable URLDisplayed as a hyperlink in the detail panel
#hidden—✓Hidden from the filter panel and the sort list, visible in detail panel
#hiddenmore——Hidden everywhere, the sort list included; useful for internal IDs
#info—Body textLong description rendered as prose in the detail panel
#ordinalOrdered categories✓Categories whose values have a natural order; numeric-prefixed values are ordered numerically
#sortquanCheckbox list✓Sorts category values by item count instead of alphabetically by default
#textlocation——Legacy address-geocoding qualifier; unsupported—use Latitude and Longitude columns instead
ℹ️

The 2026 parser also accepts #image for #img. #lat, #lon and #geometry mark the map's columns (see Map columns above); #coord is accepted for older files but places nothing.

Images & DZI tiles

How images are specified in the #img column depends on which version of SuAVE you are using. The short answer to "can I just paste in a JPEG URL?" is: yes on the hosted platform, no in SuAVE 2026.

Hosted platform (suave.sdsc.edu)

The hosted platform supports two image delivery modes. When you upload images through the New Survey dialog, the server processes them into Deep Zoom Image (DZI) pyramids server-side. Your #img column values are then the bare filenames without extension — these are image IDs that the server resolves to the right DZI tile at every zoom level.

For surveys where items already have publicly accessible image URLs (JPEG, PNG), you can also paste those URLs directly into the #img column. The platform will display them at the resolution provided. This works well for moderate-size collections where deep zoom into individual images is not needed.

ℹ️

Images are uploaded through the New Survey dialog — not through a separate tool. When creating or editing a survey at suave.sdsc.edu, the survey settings dialog lets you specify which image files to include. The server processes them into DZI pyramids automatically. There is no longer a separate image upload step via an external service.

Images go up in batches of at most 8 MB, so a slow or briefly dropped connection costs one batch, which is resent automatically, rather than the whole upload. If creating the survey then fails, the images stay on the server and pressing Create again sends only what is missing. File names without their extension become the image IDs, so they must match the #img column; non-Latin names such as Cyrillic are fine. A collection is tiled in one format: JPEG when most images are opaque, with the few transparent ones flattened onto white, PNG when most use transparency. The dialog asks only when the split is unclear.

SuAVE 2026 (self-hosted / open source)

SuAVE 2026 requires DZI pyramids for all image display. The #img column values must be image IDs, not full URLs. At load time the viewer fetches a DZC manifest file that maps each image ID to its tile location. There is no direct-URL fallback: if no DZC is configured, items without resolved tiles render as colored placeholder shapes.

# Example: #img column values for DZI delivery # The DZC URL is set on the survey (server-side or in a config file), not per-row. # Per-row values are the image IDs — bare filename stems matching the uploaded files. Image#img fossil_001 fossil_002 fossil_003

Image format and tile resolution

The DZI format (PNG or JPEG) is read from the DZC XML's Format attribute — SuAVE never hardcodes it, so mixed-format collections work correctly. At high zoom levels in Grid view, SuAVE stitches multiple DZI tiles for seamless full-resolution display. The tile level formula is min(imgMaxLevel, ceil(log2(renderSize))), with no artificial cap.

Sprite sheets (instant low-res preview)

When you upload images through SuAVE 2026's New Survey dialog, a small composite sprite sheet is generated alongside the DZI pyramid for each survey — a grid of tiny, aspect-preserved thumbnails across a few low zoom levels. The viewer loads this one small file first, so items show a real (if low-resolution) preview immediately, before the full-resolution tiles finish loading — rather than staying blank or tile-by-tile. Surveys migrated in from elsewhere with only a DZI pyramid and no sprite sheet fall back to loading tiles directly; there's no visual difference once tiles finish loading, just a slower initial appearance.

Spatial columns

SuAVE detects spatial columns by label (the part before #), not by position. The label must contain the word "latitude", "longitude", or "geometry" (case-insensitive).

Point maps (lat/lon)

Name your columns so the label contains "latitude" and "longitude" — for example Latitude#number#hidden and Longitude#number#hidden. Adding #hidden keeps them out of the filter panel while still enabling the map. Coincident points (identical lat/lon) are automatically clustered at low zoom and spread at high zoom.

Geometry maps (WKT / GeoJSON)

A column whose label contains "geometry" is treated as a WKT or GeoJSON geometry column. SuAVE renders polygons, lines, and multipolygon features on the map. The #hidden suffix is recommended to keep the raw WKT out of the filter panel.

💡

A column named country_lat#number#hidden has the label country_lat, which does not contain "latitude" and will not be detected as a spatial column. Spell out the full word.

Publishing a survey

The survey manager is at /manage. It needs an account: sign in, or use Create account on the sign-in page.

Creating a survey

New survey on the dashboard opens a two-step wizard.

Step 1, data source. Choose one:

OptionWhat happens
Upload CSV fileThe file is stored with the survey. The name field is prefilled from the file name.
Upload Corpus-DB ZIPA bibliographic network export; see Corpus-DB surveys.
Import CSV from URL onceThe server fetches the URL once and keeps its own copy.
Link to URL (re-read on every load)The survey always shows the URL's current contents; the server keeps a copy to fall back on if the URL is unreachable. Not for use with icon collections.

For Google Sheets, paste the sheet's …/export?format=csv link.

Survey name. The name is shown exactly as typed. The survey's id, which appears in its URL and file names, is derived from it: Cyrillic is transliterated (КОВРЫ22 → KOVRY22), accents are dropped, and other punctuation becomes _. Two of your surveys cannot share an id.

Step 2, configure.

  • Image definition:
    • Use icon collections: no images to upload. Items are drawn as icons whose shape and colour you assign from variables afterwards, in the editor's Icons tab (shapes, colours, or country flags).
    • Upload images to image server: select the image files (any number). Each file's name without its extension becomes its image id and must equal the item's #img value; non-Latin names are fine. Images go up in batches of at most 8 MB, so a dropped connection costs one batch, which is resent automatically. If creating the survey then fails, the images stay on the server: press Create again and only what is missing is sent. Tiles are JPEG when most images are opaque (a few transparent ones are flattened onto white) and PNG when most use transparency; the wizard asks only when the split is unclear. Progress shows uploading, then tile generation, then sprite sheets, and an email is sent when the images are ready. Upload to server chooses among the image servers your installation lists. Save full-size originals for subsequent analysis keeps each uploaded file, unchanged, beside the tiles, at /surveys/<user>_<survey>/originals/<image id>.<ext>, for notebooks and other image analysis; replacing an image in Curate replaces its original too. Without it, full-size images can still be rebuilt from the tiles later (see Curating with an AI assistant).
    • Link to existing image collection (DZC URL): use tiles generated elsewhere.
    • Create empty image repository: for surveys whose images arrive later from KoboToolbox or LimeSurvey (see Form integrations).
  • Network data (Netvis): none, a link to a network JSON file, or uploaded JSON files.
  • Image labels: always visible, on hover, or hidden.
  • Enabled views: all are on by default. The survey opens in the first enabled view (Grid unless unticked); a different starting view is set later in the editor. Netvis is added automatically when the CSV has a #netvis column.

The success screen offers Preview survey and Back to dashboard. A default About page is created, which you can fill in from the editor.

Managing surveys

The dashboard shows a card per survey with a thumbnail, title, record count and date.

  • Open launches the viewer in a new tab.
  • Edit opens the editor.
  • Delete (with confirmation) removes the survey, its CSV, About page and its own image tiles, at once; tiles are removed in the background, and the name can be reused straight away. Tiles shared with other surveys (for example by clones) are kept.

Editing a survey

The editor has seven tabs, each with its own Save, and a Preview link.

  • Info: display name (the id and URL do not change); image collection URL (DZC); preamble shown before the viewer loads; map overlay URL (Google My Maps or GeoJSON/KML) and its colour; image labels. Replace CSV data takes effect immediately: upload a new file, or fetch from a URL either once or as a live link.
  • Views: which views are enabled, and the Starting view.
  • Icons (only for surveys without their own images): assign icons from variables.
    • Shape variable and Shape collection: gender, objects, transportation, nature, science, industry, medicine, sports, technology, education, weather, food, social media, occupations, eight circle sets, or Country flags (229 flags). Values are matched to shapes automatically; for flags, country names and ISO codes such as US, UK or DRC are recognised. Numeric and date variables are split into 2 to 20 equal ranges.
    • Color variable (not with flags): 32 colours, combined with the shape variable.
  • NetVis: add network files by link or upload, or remove them. Changes apply immediately.
  • Curate: change a published survey without uploading it again. See below.
  • About: the page shown by the viewer's About button and the gallery. Read-only facts (title, date, number of variables and records) and editable Published by, Publishing date and Description / metadata / provenance. Saving replaces the page, including an automatically generated Corpus-DB About page.
  • Access: see below.

Curating a published survey

The Curate tab has four parts.

Images lists every picture in the survey's image collection, with its name, id, size in pixels (and megapixels from 1 MP up) and number of zoom levels, and a search box. SuAVE keeps the tile pyramid built from an upload, not the uploaded file, so the size shown is the original's pixel size. The zoom levels are the pyramid's: one per halving, from 1 pixel up to the full size, so ceil(log2(longest side)) + 1; a 1600×1200 picture has 12. Replace… on a picture takes a new file, shows the current and the new picture side by side with their pixel sizes and zoom levels and the new file's size (and a warning if the new one is smaller), and replaces it once confirmed; the result reads, for example, "400×300, 10 levels → 1600×1200 (1.9 MP), 12 levels". The item keeps its place, its id and its data. The new picture is tiled in the collection's own format and its cell in the low-zoom sprite sheet is redrawn, so after a reload of the viewer it shows at every zoom level. If the new file cannot be read, the old picture stays.

Replace several… takes many files at once and pairs each with the picture of the same name, ignoring the extension (IMG_0042.jpg replaces IMG_0042); upper and lower case may differ when that leaves no doubt. Files with no matching picture are listed and skipped: replacement never adds images.

Images and rows. Above the pictures, the tab checks the data's #img column against the collection, as the viewer reads it (the exact id, ignoring spaces around it):

  • Rows without their picture: values that name no picture, so the viewer shows a blank tile. Where the collection has one picture whose name differs only in case or a file extension (IMG02 for img02, img03.jpg for img03), it is offered with Use it, and Fix all takes every such match. Any other picture can be chosen by its id. Only the #img cells of those rows change; nothing else in the CSV does.
  • No row: pictures no row shows. Several rows: pictures more than one row shows, each with a count. Both filter the pictures below.
  • Empty values, N/A, and the names image_not_available and default (the original SuAVE's placeholders) mean "no image" and are only counted.
  • When the rows share a small set of pictures, as icons, flags or one picture per category do (at least ten rows with images and at most half as many different pictures), sharing and unused pictures are expected and not reported.

Clones share their source's images, and some older collections serve several surveys. When the collection is shared, the tab names the other surveys and a replacement changes the picture in all of them, after a confirmation. This is allowed only when you own every survey using the collection; otherwise the buttons are disabled. Images hosted on another server cannot be replaced here.

Variables sets how each column is read, as the original SuAVE's Tags dialog did, and rewrites only the header row of the CSV; the data are not touched.

Each column is checked against its values. Under it, a line describes them (how many, how many different, the range of numbers, a few examples), and where the data call for something else a Suggested line says what and why, with Apply. A summary above the table counts the suggestions and offers Apply all suggestions, which takes them all in one click; nothing is saved until Save. For surveys with more than 20 columns, a search box finds variables by name and Only columns with suggestions, problems or changes hides the rest (the General Social Survey's 796 columns come down to about 30). The table shows 50 variables at a time, with Previous and Next above and below it; search and the filter apply to all of them first, and Apply all suggestions and Save always cover every column, not just the page shown. The suggestions are:

  • Number when every value is a number, Ordinal when every value is a whole number from 0 to 9, and Date when every value is a date with at least a month: the rules the original SuAVE applied when it created a survey. Numbers written with a leading zero, like postal codes, are left as text.
  • Ordinal also for numbered categories such as "1. never", "2. rarely" (up to 20 of them), the way surveys like the GSS code their answers; the viewer keeps them in order and plots them at their numbers.
  • Link when every value is a web address, Multiple values when values contain |, and Long text for long, mostly different passages (a few long category labels stay categories).
  • Hidden for an empty column; for a column of 100 or more values that are all different (its list would filter nothing); and for identifiers: every value different and either a label such as "RESPONDENT ID", "case no", "SubjectID", "cruise_doi" or "ORCID", or a row count from 0 or 1. An identifier is never suggested as a Number. Hidden, it is still found by the toolbar search; choose Long text instead to search that column on its own. The column the viewer titles items with (the first plain text column when there is no #name) is never suggested Hidden or By count.
  • By count for a text column with 30 or more different values.
  • Renaming lat, lng and similar labels to Latitude and Longitude when their numbers are coordinates, so the map finds them.

Qualifiers are checked against what the viewer actually reads. A Number or Ordinal value is read from its leading number, so "89 or older" counts as 89 and "2nd important" as 2; only a value with no leading number, such as "$75+", is missing. A Date value is read if it starts with a four-digit year or the browser reads it, as with "11/21/24" or "Nov 21, 2024". An Ordinal of up to 12 ordered words ("agree", "disagree"…) is not flagged: the filter panel keeps their order, though Parallel coordinates can only plot numbers. When at least 80% of the values can be read, the qualifier is kept with a note saying how many cannot; below that it is flagged in red and another type is suggested. Types you chose that the data allow, such as Description or Item URL, are left alone.

  • Variable: the label. Renaming is allowed, with a warning: shared links, saved annotations and Jupyter or Streamlit code that use the old name will no longer match it.
  • Type: one per column; the table below lists them, as does What the choices do in the tab.
  • By count (#sortquan): values listed by how many items have them.
  • Shown: Shown, Hidden (#hidden: not in the filter panel or the sort list, still in the info panel) or Hidden everywhere (#hiddenmore: not in the info panel either).
  • Rows are numbered by column, and messages name the column too, as in Column 4 (“Title”).
  • SuAVE's own structure is shown locked: columns holding the images, item names or network data (#img, #name, #netvis), and the structural columns that stand without a label, such as #href (the page opened from an item's title) or #info (its description). As in the original SuAVE, a header that starts with # is not a variable, so such columns need no label. Change them by replacing the CSV. Qualifiers SuAVE does not offer here, such as #title, are kept as they are.
  • A qualifier needs a variable name in front of it: a bare #number, #date or #ordinal names nothing. Creating a survey or replacing its CSV with such a header is refused, with the columns listed (the original SuAVE ignored such columns); one already in a survey is hidden in the viewer and shown here as needs a label, so you can name it, for example Price#number.
  • The map takes the columns marked Latitude (#lat), Longitude (#lon) and Map shape (#geometry), whatever their labels; otherwise, as in the original SuAVE, it looks for "latitude", "longitude" and "geometry" in the labels. The review suggests the marks for coordinate or shape columns the labels would not reveal (such as lat, lng or Route), and Hidden everywhere for long shapes. #textlocation and #coord place nothing; they are kept where a file has them but not offered.
  • A survey created with Keep link has no copy of its data on the server: it reads them from the web address each time it opens, so changes made here would be lost at the next read. To edit its variables, save a copy first: Info tab → Replace CSV data → From URL, the same address, Import once (save locally), Fetch and replace.
ChoiceQualifierEffect
Textnonea list of values to tick in the filter panel
Number#numberhistogram and range slider
Date#datefiltered by date range
Long text#longsearchable from the toolbar search box, not listed in the filter panel
Link#linka link in the info panel
Item URL#hrefthe page opened when the item title is clicked
Ordinal#ordinalwhole numbers 0 to 9: a list of values in the filter panel, a numeric axis in Parallel coordinates
Multiple values#multiseveral values in one cell, separated by a vertical bar
Latitude, Longitude#lat, #lonthe map's coordinates, whatever the label
Map shape#geometryWKT or GeoJSON lines and areas for the map, whatever the label
Description#infothe item's description in the info panel

Values looks at what the cells hold and rewrites only the cells you change; the header and every other record stay as they were, byte for byte. Columns that are not filters (images, names, links, descriptions, long text, map shapes and coordinates as text) are left out. Each column lists what it found, with a suggestion where one is clear:

  • Missing-value marks such as -, NA, n/a, null or ?: the viewer shows them as a value of their own. Suggested: make them missing. Words that may be real answers ("None", "Don't know", "Refused") are suggested one at a time, never in Apply all. In columns of country or state codes, NA and DK are left alone.
  • Spelling variants that differ only in case, spaces or a final dot ("Plantae", "plantae ", "Plantae."): merged into the most frequent spelling.
  • Ranges Excel turned into dates ("Oct-49" for 10-49, "9-Jan" for 1-9), only when the column also holds real ranges.
  • Numbers inside text in a mostly numeric column ("$75+" → 75), offered one at a time.
  • Missing-data codes such as 99 or -9 far outside the other values, suggested missing; other far outliers (more than twice the largest regular value) and latitudes or longitudes out of range are shown to check, with no automatic change.

A summary counts the suggestions and Apply all takes the unambiguous ones. Any value can also be changed with Replace… or Make missing, and Show all N values lists every value of a column with its count. Columns are paged 50 at a time with a search box and Only columns with findings or changes. Nothing is written until Save. As with Variables, a survey with Keep link must be imported first.

History lists every version of the survey's data, newest first: when, by whom, rows, columns and size, and what changed in plain words, for example "Values: “-” → missing in Kingdom (4 cells)" or "Replaced the data with the file “birds.csv” · 60 → 58 rows". Every change to the CSV is recorded, whether made in Curate (Variables, Values, image references) or by replacing the CSV in the Info tab. The data a survey had before its first recorded change are kept as version 1.

  • Compare with current shows which columns were renamed or retyped, how many rows were added, removed or changed, and the first 20 changed cells.
  • Download saves that version's CSV, exactly as it was.
  • Restore… makes that version the current data again, after a confirmation. The data it replaces become a version too, so a restore can be undone. Icon shapes and colours (Icons tab) follow their variable if the restored header names it differently. A survey created with Keep link cannot be restored here, as its data come from the link.
  • The newest 5 earlier versions are always kept, and up to 20 for 30 days. When a copy is removed its entry stays in the list ("copy removed").
  • Provenance record (JSON-LD) downloads the history in the W3C PROV standard (PROV-O): each version is a prov:Entity, each change a prov:Activity that used the previous version and generated the next, by a prov:Person. The sentences in the list are written from this record.

Curating with an AI assistant

An AI assistant such as Claude Code or OpenAI's Codex can do the Curate tab's work in conversation: review a survey's variables, values, images and history, explain what it finds, and make the changes you approve. It works through SuAVE's MCP server (the Model Context Protocol, the standard way assistants use outside tools) with an access token instead of your password.

In the gallery, AI assistant opens the setup. The simplest way is to connect by address, https://suave.sdsc.edu/manage/api/mcp, with nothing to install:

  • Claude in the browser or Claude Desktop: Settings → Connectors → Add custom connector, name it SuAVE, paste the address, then Connect. SuAVE's sign-in page opens: sign in and Allow. Claude Desktop shares the browser's connectors. On a Team or Enterprise plan, an owner may have to add the connector first.
  • Claude Code: claude mcp add --transport http suave -s user https://suave.sdsc.edu/manage/api/mcp, then /mcp in Claude Code, choose suave, Authenticate.
  • Codex: codex mcp add suave --url https://suave.sdsc.edu/manage/api/mcp, then codex mcp login suave.

The sign-in page names the app and the site it will return to; continue only if you started the connection there. The connection appears in the panel's list ("Claude (connected app)") and can be revoked there. It renews itself for 90 days of use.

The SuAVE MCP server can also run on your own computer, with an access token instead of the sign-in (the panel's "Or run SuAVE's MCP server on your computer"):

  1. Download the SuAVE MCP server (one file, which needs Node.js 18 or later) and put it in a suave-mcp folder in your home folder: the panel gives the terminal command, mkdir -p ~/suave-mcp && mv ~/Downloads/suave-mcp.js ~/suave-mcp/.
  2. Make a token for each assistant, named after it ("Claude Code on my laptop"), so History shows which one made a change. Copy it at once: it is shown only then.
  3. Run the panel's command for your assistant in a terminal, claude mcp add suave … for Claude Code or codex mcp add suave … for Codex; it already contains the token. For Claude Desktop, put the panel's settings into Settings → Developer → Edit Config, with the two paths filled in. If SuAVE was added before, remove it first (claude mcp remove suave -s user, codex mcp remove suave).

Then start the assistant again and ask, for example, "List my SuAVE surveys" or "Review the values in my survey Birds and suggest fixes". The assistant can list your surveys, read each one's variables (with a profile of every column), problems in the values, images against rows, and versions, and compare a version with the current data. It changes qualifiers and labels, values, image references, or restores a version, only with an explicit list of changes.

It can also look at the pictures. get_survey_definition says what a survey is made of: rows and columns, whether its images are on this server, their format and sizes (smallest, median and largest side, how many over 2000 px), zoom levels, and how many full-size images are kept. get_image shows it one picture, whole and scaled down or a region at full resolution (up to 2048 px across), so it can examine detail such as brushwork or handwriting. reconstruct_full_images_from_dzi writes the full-size image of chosen images, or all, rebuilt from their zoom tiles (the top zoom level has each image's full pixel size), into the same originals/ folder; it runs in the background, never overwrites an uploaded original, and list_full_images gives the files' addresses for a notebook. Larger analyses, such as comparing brushstrokes across thousands of paintings, belong in a notebook that reads those files.

A connected app or a token can only curate your surveys: it cannot delete them, change their settings, upload files, sign in or make other tokens, and administrator rights never pass through it. Every change it makes is saved in the survey's History as made by that token on your behalf ("you via “Claude Code on my laptop”"), so any change can be compared and restored. Tokens expire after a year; Revoke… stops one at once. Never paste a token into a chat or a shared document; if one was, revoke it and make another.

Showing images, icons or flags

You haveChooseNotes
Photos, scans or artworkUpload images#img holds each image's file name without extension
A tile collection made elsewhereLink to existing DZCPaste the .dzc URL
No images, but categoriesUse icon collectionsThen set Shape/Color variables in Icons
CountriesUse icon collections → Country flagsValues may be names or ISO codes
A Corpus-DB exportUpload Corpus-DB ZIPFlags are set automatically from the authors' countries
Images arriving from a formCreate empty repositorySee KoboToolbox and LimeSurvey

Cloning a survey

You can copy any survey you can open into your own account: public surveys, your own, and private surveys that list your email. Use Clone to my gallery in the viewer's toolbar (or in the Jupyter dialog). After signing in you are asked for a name, "<name> (clone)" by default, up to 200 characters, and the copy opens in the editor.

  • Copied: the CSV (a linked survey becomes a local copy, no longer live), the About page, views and starting view, label setting, preamble, map overlay, icon assignments and Streamlit settings.
  • Shared, not copied: the image tiles and network files. Deleting the original does not remove tiles a clone still uses.
  • Reset: the clone is visible in your gallery and has no access restrictions.

Access and the public gallery

  • Hide from public gallery removes the survey from your gallery; its link still works.
  • Restrict access by email list (one address per line, or comma-separated) makes the survey private: only you, and signed-in accounts whose email is on the list, can open it; others see a sign-in prompt. Private surveys are also kept out of the public REST API, Jupyter and Streamlit. Image tiles and the About page are not restricted.
  • Every account has a public gallery at /gallery/<username> listing its visible surveys, with Open and About buttons.

Loading a CSV without an account

SuAVE 2026 can open any CSV directly from the landing page: drag and drop a file, or paste a URL (Google Sheets links are converted automatically). No account is needed, and a dropped file never leaves your browser. Icon-based and text-only surveys work fully this way; image collections need a configured tile (DZC) URL.

Corpus-DB bibliographic networks

Corpus-DB is a free service that builds global bibliographic networks from OpenAlex: you describe a network by a thematic publication scope or by a list of authors, refine it by keywords, countries, institutions and publication years, and download it as a dataset. Imported into SuAVE, it becomes a survey of authors that you can filter, map, compare, and explore as a co-authorship network, following how it grows over time.

The whole workflow has four stages: generate, clean, (optionally) merge, and publish.

1. Generate a dataset

Go to corpus-db.sdsc.edu and sign in with your SuAVE username and password. Choose Search, check that the email at the top of the page is yours (results are sent there), and give the project a name. Then choose one of two search types.

Scope (important terms). Use this to build a network around a topic.

  • Scope: the words and phrases publications must contain. Follow the on-page instructions closely, including plural forms of words.
  • Allow relaxed search: match the scope words in any order. Leave unticked to require the scope exactly as written.
  • Exclusions: words that must not appear.
  • Keywords: words to tag authors with when they appear in a title or abstract, with an option to exclude publications that contain none of them.
  • Starting year / Ending year: limit the publication years.
  • Institutions and Collaborating institutions: comma-separated, listing every form of each name (for example UC San Diego, UCSD, University of California San Diego) so that nothing relevant is missed.
  • Too many co-authors: optionally leave out authors with more than a threshold of co-authors (25 by default), which keeps very large consortium papers from swamping the network.

Author OpenAlex ID. Use this when you already know the people: enter their OpenAlex author ids, comma-separated. Options are the same as above (keywords, years, institutions, co-author threshold), plus whether to include external authors in general, or only those within the authors' institutions.

Press Submit. Depending on the size of the search, an email arrives within seconds, minutes or, for very large ones, hours. Follow its link to corpus-db.sdsc.edu/collect, enter the code from the email, and the dataset downloads. If the first submit shows an error, open a new tab, return to Corpus-DB and sign in again.

The download contains three files: the authors CSV, a *-publications.csv, and the network JSON. If they arrive as a folder rather than a ZIP, put all three into one ZIP for SuAVE.

2. Clean it

Bibliographic data repeats people under slightly different names. Clean it before publishing:

  1. Create a SuAVE survey from the ZIP (see stage 4); Corpus-DB's curation works on SuAVE surveys.
  2. In Corpus-DB, open Curation, select the dataset and choose Select or Upload Survey.
  3. Delete rows whose names are not real names, such as "et al.", "anonymous" and their variants. Click a row (or double-click) to edit it; Save Changes, Delete Row and Close are at the bottom of the page.
  4. Remove Non-Latin (ISO/IEC 8859-1) cleans characters that would otherwise split one name into several.
  5. Merge duplicate people: Match Duplicated Name (same name) or Match Duplicated Name/Affil (same name and affiliation, more selective). Review the candidates at the bottom of the page, select them (or Select All) and Merge.
  6. Fuzzy Match Name finds near-identical names. For each group, choose Select All and Merge, Merge Selected or Skip to Next; "Choose item into which others will be merged" decides which row survives (Auto by default). Repeat until no more matches are offered.
  7. Save Results downloads the cleaned ZIP. Upload it to SuAVE as a new survey, and delete the uncleaned one.

For harder cases, the cleaned authors CSV can be curated further in OpenRefine: load it, cluster the Name column (Edit cells → Cluster and edit), merge clusters so that one person has one spelling (first name then last name, spelled out where possible), try other clustering methods including an n-gram size of 1, remove rows such as "Anonymous" or "Author" with a text facet, and finally add a last-name column (value.split(" ")[-1]) and review clusters by last name, merging only where the other fields agree. Save as UTF-8.

3. Merge datasets (optional)

The CorpusDB Merge tool combines several datasets or networks. Its operations are: import CSV, merge CSVs (on a key such as the OpenAlex id, choosing whether overlapping rows are overwritten), conjoin networks, add a facet (column) to a CSV, edit or delete facets, build a network from chosen facets, and export the result as CSV, JSON or a ZIP for SuAVE. Steps are connected visually, by dragging from one block's tab to the next; save the network before leaving. A typical use is to conjoin two networks over the same people, for example co-authorship and shared research concepts, so that each author node links both to collaborators and to topics. The tool's Help button has a video tutorial and worked examples.

4. Publish in SuAVE

In the survey manager choose New survey → Upload Corpus-DB ZIP, select the ZIP, name the survey and create it. What the import does:

  • Data: the authors CSV (the first .csv in the ZIP that is not *-publications.csv) becomes the survey. Each row is one author, with columns such as the OpenAlex id, affiliation, city, region and country (with coordinates), OpenAlex concepts, total publications, first and latest publication dates, citedness, H index, the search scope and keywords, collaborators in scope, publication dates, a Show Publications cell, #img (the author's country code) and #netvis (the author's key in the network).
  • Network: every JSON file in the ZIP is attached, so the Netvis view shows the co-authorship network, with each author's node linked to their row.
  • Images: authors appear as country flags; #img holds ISO country codes (un where unknown), matched to SuAVE's shared flag collection.
  • About page: generated automatically, with the number of publications (from *-publications.csv), the number of authors, the span of publication years, the search scope, its keywords and the most prominent OpenAlex concepts. Replace it any time from the editor's About tab.

Show Publications. Each author's Show Publications button, in the info panel, opens a page listing that author's publications within the scope, retrieved live from Corpus-DB.

Author photos instead of flags. A notebook in the SuAVE documentation finds candidate photos of authors from their name, affiliation, city and country. To use photos, set each author's #img value to their photo's file name (without extension), and either create the survey with Upload images in step 2 of the wizard, selecting the photos, or attach an existing tile collection from the editor's Info tab.

Exploring. Map the authors and colour them by H index or citedness; compare countries and concepts in Crosstab or Heatmap; follow first and latest publication dates in Timeline; and in Netvis, move the year range to watch the network grow, trace shortest paths between two authors, and open node statistics to find brokers and central figures.

Step-by-step guides with screenshots are on the SuAVE documentation site: generating a dataset, cleaning it, curating names in OpenRefine, the Merge tool, conjoining networks and author photos.

Sample datasets

The SuAVE team maintains a curated set of sample datasets for workshops and experimentation, covering three broad types.

Surveys without images

These use icon specifications (shape and color driven by variable values) rather than photographs. The EarthCube Member Survey and various NSF award databases fall into this category. They are the simplest starting point: prepare a CSV, annotate the columns, upload — done.

Collections with limited icon sets

Surveys like the Observing Systems Explorer and 2018 SDG Indicators dataset use small, bounded image sets that are easy to process. Good for learning the image upload workflow without managing thousands of files.

High-resolution image collections

Examples include the Picasso Paintings survey, the USGS Earth As Art collection, the Van Gogh paintings collection, BGS Macrofossils, and wetland soil samples. These demonstrate the full DZI tile pipeline and deep-zoom interaction.

Download the sample datasets

Each dataset is a public Google Drive folder with a README describing the data and its source, the CSV, and, where the survey has them, the images, ready to publish in SuAVE. The same list is kept in the SuAVE sample datasets document.

DatasetShows how to publish
Picasso paintingsHigh-resolution images
Earth as Art (USGS)High-resolution images
EarthCube member survey 2013No images: icons from variables
Earth Observing SystemsA small set of icons and logos
Wetland samples, Lower MekongHigh-resolution images
San Diego vacant lotsHigh-resolution images
2018 SDG IndicatorsA small set of icons, and a map
Jordan foodImages

Published versions of many demo surveys can be browsed in the 100+ Surveys gallery.

Practice with live data

A good first exercise: go to the NSF Awards search, run any keyword search, export the results as CSV, add a #number qualifier to the Amount column and a #date qualifier to the start date, then upload. You get a funded-project browser with a histogram and date filter in under ten minutes, no images required.

Available views

Every view shows the same filtered set of items; switching views keeps the filters (and closes the info panel). How a view treats items with no value for its variable is controlled by Settings → Display missing values (off by default): when off, such items are left out of grouped views; when on, they form an N/A group, placed last.

Grid

Every filtered item as a square tile, in the toolbar's sort order. At zoom 1 the whole set fits the canvas; zooming keeps the column count and enlarges the tiles, up to the resolution of the source images.

  • Wheel or pinch zooms around the pointer; drag pans.
  • Click a tile to open the info panel. A small tile (under about 112 px) is also zoomed to about 180 px and centred.
  • Double-click opens the lightbox: a large image with the fields chosen in Settings → Lightbox fields. In the lightbox, ← and → step through items and Esc closes it.
  • Keyboard (click the grid first): arrow keys move a focus ring, Enter or Space opens the info panel, Esc closes it.
  • Small tiles come from a sprite sheet and switch to deep-zoom images as you zoom in; items without an image show an icon or a coloured placeholder. Very large collections draw progressively.
  • Tile labels (Settings) show each item's name under tiles at least 48 px wide.

Bucket

A histogram of image stacks: one column per group, tiles stacked from the bottom, column height proportional to the count.

  • Variable: any variable except long text, links, images, coordinates and the name column.
  • Group by (not shown for numbers):
    • Alphabetical ranges (initial): up to 10 contiguous ranges such as Albania–Chile; if every value is a number, numeric ranges instead.
    • Alphabetical (individual): one column per value.
    • Natural order: only when SuAVE detects an ordered sequence such as Low / Medium / High.
    • Largest first: columns by count.
  • Other: in the last three modes, beyond 10 columns the largest 9 are kept and the rest are combined into Other, meaning "values not shown", not missing data. Hide Other / Show Other toggles it.
  • Numbers are split into 10 equal-width bins over the whole dataset's range, so the bins line up with the filter slider. Items with no number are never shown here.
  • Column labels show the value or range, its count and its share of the filtered items. While filters are active a contribution figure is added (filtered share minus unfiltered share), red above +10 points, blue below −10, green otherwise. Clicking a label shows its full text.
  • Contributions opens a table under the canvas (only while filters are active) with filtered and unfiltered counts and shares for every column, and one "Contrib" column per active filter: how much that one filter changes the column's share.
  • Wheel zooms around the tile under the pointer; pinch zooms; drag pans; click a tile opens the info panel.

Bars

One vertical column of tiles per group, side by side. Each column scrolls on its own and the row of columns scrolls sideways, so every group can be browsed in full.

  • Variable: as in Bucket.
  • Mode: Alphabetical ranges, Alphabetical (individual), Natural order (when detected), Largest first. There is no Other column: in the non-range modes every value gets a column.
  • Numbers and dates in Alphabetical ranges use "nice" power-of-ten bins, about ten at most, labelled with K/M/B abbreviations unless that would make two bins look alike.
  • Zoom with the slider or a pinch (no wheel zoom on the canvas). Drag the background to scroll sideways, or inside a column to scroll it.
  • Click a tile to open the info panel; the panel's arrows then step through that column only. Hovering a tile shows its value.

Crosstab

A two-way table: one variable down the rows, another across the columns. Each cell shows thumbnails of its items under a count overlay coloured by size.

  • Rows and Cols: any groupable variable.
  • Group by: the same four modes as Bucket, applied to both axes.
  • Counts cycles the count overlay: solid, faded, off.
  • Hide/Show Other appears when an Other group exists.
  • Numbers and dates are split into up to 10 equal-width ranges over the filtered data; a variable with one value throughout is a single column named by that value. Numeric rows run upward, smallest at the bottom. Text with many values is grouped into alphabetical ranges, numeric codes into numeric ranges.
  • Cells show the count and the share of the row. Up to 50 thumbnails per cell, sized by the cell's relative count.
  • Header click: shows exactly the items that header counts, replacing any selection on that variable. A text header selects all of its values, a number or date header the span of its values. The axis then regroups within the selection, so a range header splits into finer ones you can click in turn. Clicking a header that already is the whole selection clears it. The active header is outlined.
  • Opening a cell: with Counts on, click the cell; with Counts off, use its ⊞ button. The detail view behaves like Grid (wheel, drag, click), and "← Overview" returns. The info panel's arrows stay within the cell.
  • Wheel or pinch on the overview enlarges the cells, up to 20 times.

Map

Items as points, lines or areas on an OpenStreetMap base map.

  • Which columns are used: a column whose label (the part before #) contains "latitude" and one containing "longitude" give points; a label containing "geometry" gives WKT lines and polygons. #hidden columns work, e.g. Latitude#number#hidden. Without such columns the view says so.
  • Color by: any variable.
    • Text and other categories: a palette of up to 12 legend entries; more values are grouped into ranges. A multi-value cell uses its first value.
    • Numbers and dates: a colormap (Cool-Warm by default, or Plasma, Viridis, Blues, Greens), classified by Quantiles (default), Equal intervals or Natural breaks (Jenks), into 2 to 20 classes (default 5).
    • Items with no value are always drawn, in neutral grey. Their N/A legend entry, with a count, appears when Display missing values is on.
  • ⌂ Fit all zooms to the data. Legend toggles the legend. Cluster (points only, on by default) groups nearby points; with it off, points sharing a location fan out around it. Tooltips toggles hover tooltips.
  • ⚙ Settings (for numeric colouring, or when there are areas or lines): area Opacity (default 85%), line Width (1 to 8 px), Colormap, Classification, Classes, and the overlay colour.
  • Hover (desktop) shows the name and colour value. Click opens the info panel; double-click opens it at once. The selected item gets a frame, and the map pans to it if it is off screen.
  • Overlay: a survey's map overlay URL adds a reference layer: a Google My Maps link (its KML is fetched), or any KML or GeoJSON URL.

Table

A spreadsheet of the filtered items: a thumbnail column (when the survey has images), then every visible column in CSV order. Hidden columns and coordinates are left out. Cells show up to five lines and scroll inside.

  • Click a header to sort; click again to reverse. This sort is local to the table; the starting order is the toolbar's.
  • Drag pans both ways; click a row to open the info panel. Rows are rendered only as they scroll into view, so large tables stay fast.

List

A compact list in the toolbar's sort order. Each row shows a thumbnail, the name and one line of description: the first #info column, otherwise the first #long column. Click a row to open the info panel.

Network

A graph built on the fly from the survey's own columns.

  • Connect: the variable whose values become one set of nodes (indigo). To: the variable whose values become the other set (green). With: the variable whose shared values create the edges.
    • Connect = To (co-occurrence): two values are linked when items sharing a With value carry them. Groups are capped at their first 50 items.
    • Connect ≠ To (bipartite): each item links its Connect values to its To values, labelled with the With value.
  • Style: Force directed (default), Circle, Concentric, Breadth first, Grid, Random.
  • Edge width grows with the number of items behind the edge. A Min weight slider hides weak edges; Node connections min/max hides nodes outside a range of total connection weight.
  • Large graphs: above 250 nodes the view keeps the 250 most densely connected (by k-core) and says so in a banner.
  • Click a node to highlight its neighbours and list its connections (with the linking value); click it again, or the background, to restore. Click an edge to highlight it. Right-click (or Ctrl+click) a node to open the info panel for an item with that value. Clicking does not filter.
  • Items with no value are left out; multi-value cells are split on |.

Netvis

A pre-computed network supplied with the survey as JSON (typically a co-authorship network from a Corpus-DB export). The tab appears whenever the survey has a #netvis column.

  • How items connect to the network: an item's #netvis value is a key into the JSON's data. Each entry lists the item's connections (each with a name, category, optional weight, year list, icon and links). The item itself is the node named by its #name value. The graph is rebuilt from the filtered items, so filters shrink it.
  • Node icons: Country flags (default) use each entry's flag icon; Survey images use the items' own thumbnails, falling back to the flag.
  • Toolbar: Controls, Restore (after isolating or tracing a path), zoom in/out, pause/resume the layout, zoom to fit, a dataset selector when there are several network files, and node/edge counts.
  • Controls panel:
    • Year range (when the network has years): hides nodes outside the range.
    • Hops from selected node: 1 to 6, or All.
    • Min edge weight: hides weaker edges.
    • Size nodes by: collaboration (degree), any column the network declares, or uniform; Node scale 0.5 to 5 times.
    • Graph statistics: average clustering and fragmentation.
    • Networks: one checkbox per network when the file holds several.
  • Click a node to isolate its neighbourhood. Shift+click two nodes to trace the shortest path between them. Alt/Option+click (or Cmd+click) opens node statistics: betweenness, prestige, clustering, square clustering and key-player fragmentation. Right-click (or Ctrl+click) opens the info panel. Double-click the background to zoom in; click it to restore.
  • Network file format (v2): ``json { "config": { "version": 2, "view": { "<network>": { "slider_label": "Year", "node_scaling_options": ["H Index"], "max_weight": 10, "connection_title": "Co-authors of %s" } } }, "data": { "<key from the #netvis column>": [ { "category": "Author", "network": "<network>", "slider_variable": [2019, 2021], "icon": "assets/flags/US.png", "links": ["https://openalex.org/…"], "connections": [ { "name": "Jane Doe", "category": "Author", "weight": 3, "slider_variable": [2020], "icon": "assets/flags/FR.png" } ] } ] } } ` Files without config.version` are read as the older v1 layout (one entry per key).

Heatmap

A colour matrix of two variables.

  • Rows, Columns: text, multi-value, number, date or ordinal variables, at most 40 labels per axis.
  • Connect via (optional): the tooltip lists the top values of a third variable in each cell.
  • Sort: by count (default) or A to Z. Values: counts, % of row, or % of column. Colors: eight scales, Blue→Red by default, with a draggable legend.
  • Hover for counts and percentages. Click a cell to filter to its row and column values (numbers and dates by their value span).
  • Missing values follow Display missing values. Export as PNG.

Scatter

Two numeric variables against each other.

  • X, Y: numbers. Color: a category (top 12 values, plus N/A and Other). Size: a number (4 to 20 px).
  • Points: dots, or the items' images when the survey has them (Icon size 12 to 64 px); a warning appears above 800 images.
  • Zoom: wheel or pinch zooms both axes around the pointer, and dragging pans. Reset zoom appears while zoomed. The zoom survives filter changes and resets when X or Y changes.
  • Rows without both X and Y are left out. Click a point to open the info panel, zoomed or not. Export as PNG.

Timeline

Item counts over time, as stacked bars.

  • Date field, optional Group by (top 12 values plus Other), and Period: Auto, Decade, Year, Month, Week or Day. Auto picks decades for spans over 15 years, years over 2, months over 60 days, weeks over 14 days, otherwise days.
  • Zoom: wheel or pinch zooms along the time axis, dragging pans, and the slider under the axis shows and moves the visible span. With Period on Auto, zooming in switches to finer periods by the same rule, applied to the visible span: a decade splits into years, then months, weeks and days, and the count axis rescales to the bars in view. Shift+wheel, or the slider at the left, stretches the count axis so thin segments of a grouped bar can be read. Reset zoom appears while zoomed. The zoom resets when the date field or Period changes.
  • Click a bar to filter to that period. Undated items are left out. Export as PNG.

Parallel coordinates

One line per item across parallel numeric axes.

  • Axes: up to 8 number or ordinal variables (the first 4 pre-selected), shown as removable chips with an "add axis" selector. Color: a category.
  • Drag along an axis to highlight a range; lines outside it are dimmed. This highlights only; it does not filter. Export as PNG.

Sankey

Flows between categories.

  • From, optional Via, and To variables; Min flow hides thin links. Each column keeps its top 30 values.
  • Click a link to filter to both its ends; click a node to filter to that value; drag nodes to rearrange. Export as PNG.

Info panel, toolbar & settings

Info panel

Click an item in any view to open it.

  • Navigation: arrows step through the current items (or, in Bars and a Crosstab cell, through that group), and the view scrolls to follow. The counter shows the position.
  • Title: the #name value, a link when the survey has an #href column.
  • Image: a zoomable view of the item's image.
  • Body text: any #info column, as formatted text (HTML is cleaned before display).
  • Fields: every other visible column. Clicking a text, ordinal or multi-value entry filters to that value; multi-value entries are separate pills; #link values are links. A cell containing getPublication({...}) becomes a Show Publications button (see Corpus-DB surveys).
  • Notes: a private notes field per item, saved in this browser.
  • Layout: Settings → Info panel chooses Shrink canvas or Overlay. On phones the panel is a bottom sheet.

On phones these move to a bar at the bottom.

Advanced Search

A multi-condition filter. Match All (AND) or Any (OR). Text conditions: contains, exactly equals, does not contain; numbers: between; dates: from/to. In All mode the conditions become ordinary filters with chips; in Any mode the matching set is applied directly, without chips, and is not part of share links.

Share / copy link

Builds a link that reproduces the current view, sort, filters, search and selected item, with a Copy button and shortcuts for Facebook, LinkedIn, WhatsApp, X and email. Calculated columns are not included.

Clone to my gallery

Shown for surveys stored on the server. Opens the survey manager with a copy of this survey ready to be made (see Cloning a survey).

Download filtered CSV

Saves the filtered items as {survey}_filtered.csv, with the original columns in order, dates as YYYY-MM-DD and missing values as empty cells.

Calculate new variable

Builds a new numeric column from existing ones. Each row is a number column or a constant, optionally transformed (Σ sum over all rows, −x, √x); rows are joined by + − × ÷ and evaluated left to right, without operator precedence. A missing input, division by zero or the square root of a negative gives a missing result. The column exists for this session only.

Annotate / comment on pattern

Saves a comment with a snapshot of the current filters and search as a named pattern in this browser. Saved patterns can be restored (filters and search are re-applied) or deleted. Unlike per-item notes, a pattern describes a subset.

Settings

Changes apply when you press Apply; all but the survey name are remembered in this browser.

SettingWhat it does
Survey nameDisplay name for this session only
Display missing valuesShow items with no value as an N/A group in grouped views and the map legend; the hint counts them for the current variable
Tile labelsAlways, Hover or Never, and the Label column
Default sort orderAscending or descending
Info panelShrink canvas or Overlay
Visible variables in filter panelMove variables between shown and hidden lists
Lightbox fieldsUp to 8 fields shown in the Grid lightbox

About this survey

Shows the survey's About page in a dialog. Escape closes it.

Save chart as PNG

For Heatmap, Scatter, Timeline, Parallel and Sankey: saves the chart at twice screen resolution.

Enter fullscreen

Fills the screen; Escape leaves.

Interface language

Add ?lang= to the viewer URL to choose a language: ar_AR, es_MX, ka_GE, kk_KZ, ru_RU, tr_TR, zh_CN or zh_TW. English is the default, and anything not yet translated stays in English.

Filtering & search

The filter panel lists every facet column as a collapsible accordion. Checking or unchecking values updates all views in real time — no page reload. Multiple selections within one facet combine as OR; selections across different facets combine as AND.

Search

The toolbar search box searches across all columns by default. When a #long column is selected in the scope dropdown, search narrows to that field. Within each facet accordion, a separate "Filter values…" input filters the checkbox list display — it does not affect item counts.

Advanced search

The search icon in the toolbar opens the multi-condition filter builder. Match all conditions (AND) or any (OR). Text conditions are contains, exactly equals and does not contain; numbers take a range, dates a from/to. See Info panel, toolbar & settings.

Filter breadcrumbs

A row of colored chips below the toolbar summarizes all active filters. Click × on any chip to remove that filter. A Clear All button in the filter panel header removes everything at once.

Annotation mode

The annotation icon saves the current filter state — every active checkbox, range, and search term — as a named pattern in the browser's localStorage. Restore any saved pattern with a single click. Useful for teaching, where students document their analytical steps.

Sharing & export

Share link

The share icon builds a link that reproduces the current view, sort order, filters, search and selected item, with a Copy button and shortcuts for Facebook, LinkedIn, WhatsApp, X and email. Calculated columns are not included.

Download filtered CSV

The download icon exports the currently visible items as a CSV file — only rows that pass all active filters are included. Reserved variable names and qualifiers are preserved in the header.

Chart export

For chart views (Heatmap, Scatter, Timeline, Parallel, Sankey), a camera icon in the toolbar downloads the current chart as a high-resolution PNG.

Survey config files

A survey config file is a plain JSON document that bundles all the information needed to open a survey without a backend account. It is useful for self-hosted deployments, for sharing a survey with a specific pre-selected view or overlay, or for attaching Streamlit apps to a CSV that lives on an external URL.

// Example survey config file (save as mysurvey.json) { "csv": "https://example.com/data.csv", "dzc": "https://example.com/tiles/content.dzc", "fullname": "My Survey Title", "views": ["grid", "bucket", "bars", "map", "table"], "map_overlay_link": "https://www.google.com/maps/d/...", "show_image_labels": "always", "preamble": "<p>Welcome to this survey.</p>" }

Only csv is required. All other fields are optional. Load a config file with the URL parameter ?surveyconfig=URL. Google Sheets share links in the csv field are automatically converted to their export equivalent.

URL parameters

SuAVE 2026 accepts several URL parameters for opening a specific survey directly.

ParameterPurposeExample
?survey=filename.csvLoad a CSV from the server's surveys directory?survey=suavelocal_BGS_Macrofossils.csv
?user=X&file=YOriginal SuAVE URL style; equivalent to ?survey=X_Y.csv?user=suavelocal&file=BGS_Macrofossils
?csv_url=URLLoad any CSV from an arbitrary public URL?csv_url=https://…/data.csv
?surveyconfig=URLLoad a JSON config file (CSV + DZC + views + metadata)?surveyconfig=/configs/my.json
?lang=CODEOpen the interface in another language (see Interface languages)?lang=ru_RU
?skip_preamble=trueSuppress the intro modal on surveys that have one

Interface languages

Add ?lang=CODE to any survey URL and SuAVE opens in that language. Only the interface is affected: labels, buttons, menus, dropdowns, and the inline help lines under each view. Survey data is never translated, so variable names and values appear exactly as they are in the CSV.

Currently available:

LanguageCodeNative name
English (default)en_USEnglish
Arabicar_ARالعربية
Chinese (Simplified)zh_CN简体中文
Chinese (Traditional)zh_TW繁體中文
Georgianka_GEქართული
Kazakhkk_KZҚазақша
Russianru_RUРусский
Spanishes_MXEspañol
Turkishtr_TRTürkçe

A complete example:

https://suave.sdsc.edu/view?survey=suavelocal_BGS_Macrofossils.csv&lang=kk_KZ

The parameter combines with every other URL parameter, and it survives the share link, so a link you copy while working in Kazakh reopens in Kazakh for whoever you send it to. Leaving the parameter off gives you English, and so does an unrecognized code.

Adding a language

The list grows as people ask for it. If you would like to work with SuAVE in a language that is not here, write to izaslavsky@ucsd.edu and we will send you the file to fill in. A translation is one JSON file of short interface strings, with the English on the left and your language on the right, and it takes an afternoon rather than a project.

Self-hosted deployments can add one directly. Drop the finished file into the frontend's public/ directory under its language code, say fr_FR.json, and it is served by name the first time someone passes ?lang=fr_FR. Nothing needs to be registered and the application does not need rebuilding.

Jupyter & Streamlit integration

SuAVE hands the survey you are looking at, filters included, to Python: to a Jupyter notebook for analysis, to a Streamlit app, or to any script through the REST API. Results can come back as a new survey.

Turning them on

The Jupyter and Streamlit buttons appear beside the view tabs when two things are true: the survey has jupyter or streamlit ticked in the editor's Views tab, and the server's defaults.json lists at least one Jupyter or Streamlit server. The server list is set by the administrator for the whole installation; authors only switch the buttons on or off per survey. Private surveys are not available to these integrations.

Jupyter

The Jupyter dialog has up to three tabs.

  • JupyterHub (when servers are configured): choose a server and Connect. The notebook opens with the survey passed in its URL: the survey id and CSV, its image collection, the survey URL, the API address, the enabled and active views, and the current filter state.
  • Binder: give a GitHub repository, branch and notebook (by default izaslavsky/suave-notebooks, main, SuAVEDispatch.ipynb; your choice is remembered). SuAVE stores the current survey and filters for 30 minutes under a one-time token, starts the repository on mybinder.org, and the notebook picks the parameters up automatically. The repository must include SuAVE's small receiver (suave/receiver.py), as suave-notebooks does.
  • Colab: opens the notebook in Google Colab and shows SUAVE_TOKEN and SUAVE_HOST, with Copy buttons, to paste into the notebook's first cell. The token also lasts 30 minutes.

The dialog also offers Clone to my gallery.

The suave-notebooks collection. SuAVEDispatch.ipynb is a menu of ready-made analyses that run on the survey you sent:

GroupNotebooks
Statisticsdescriptive statistics, contingency tables, factor contributions, supervised classification and regression, PCA, clustering, outlier detection
Arithmetic and wranglingderived variables, variable transforms
Spatialgeographically weighted regression, exploratory spatial analysis, aggregate maps
Networksnetwork building and metrics for Netvis surveys
Imagescolour statistics, image classification; AI captioning (BLIP), object detection (DETR), image clustering (CLIP)
Textsentiment, zero-shot classification, geocoding, topic modelling, entity linking to Wikidata, OpenAlex enrichment

Image notebooks need the survey's full-size images: uploaded originals, or full-size images rebuilt from the tiles (see Curating with an AI assistant), at /surveys/<collection>/originals/; network notebooks need a #netvis column.

Publishing back. A notebook can create a new SuAVE survey from its results. It asks for your SuAVE username and password and creates the survey in your account.

Streamlit

The Streamlit dialog lists the installation's Streamlit servers; Connect opens the chosen app with the survey (user, CSV, image collection, survey URL, API address and filter state) in its URL. The standard launcher, suave-launcher, forwards to:

  • suave-arithmetic: new numeric variables from arithmetic on existing ones;
  • suave-spatial-gwr: geographically weighted regression, residual maps and Moran's I;
  • suave-factor-contribution: tests rules such as "if A and B then C" against the data.

The arithmetic and GWR apps can publish their result as a new survey, after you sign in with your SuAVE account.

Survey integrations

SuAVE connects directly to KoboToolbox and LimeSurvey through their REST APIs. Once configured, new survey submissions flow into SuAVE automatically — no CSV exports, no scheduled imports, no reformatting.

KoboToolbox

KoboToolbox is the standard tool for humanitarian and field research surveys. It supports GPS location capture, photo uploads, offline data collection, and branching logic — making it the right choice for ecology surveys, archaeological fieldwork, urban assessments, and any study where respondents are not at a desk.

The integration uses KoboToolbox's REST service feature. To connect a Kobo form to SuAVE:

  1. In your KoboToolbox account, open the form you want to connect. Go to Settings → REST Services → Add Service.
  2. Set the endpoint URL to your SuAVE instance's Kobo ingest endpoint: https://suave.sdsc.edu/kobo/submit?user=YOUR_USERNAME&file=YOUR_SURVEY.csv
  3. Add your SuAVE API key in the Authorization header field.
  4. Save. From this point, every new Kobo submission triggers a POST to SuAVE, which appends the row and refreshes connected viewers.
📱

GPS and images map automatically. Kobo geopoint fields are split into columns whose names contain Latitude and Longitude, so map view activates without additional qualifiers. Photo fields populate the reserved #img variable, which supplies the image grid.

LimeSurvey

LimeSurvey is widely used for course-integrated research and institutional questionnaires. The SuAVE–LimeSurvey bridge uses LimeSurvey's RemoteControl API to pull responses on a configurable schedule (or on demand from the SuAVE authoring panel).

  1. In LimeSurvey, enable the RemoteControl 2 API under Global Settings → Interfaces and create an API user.
  2. In SuAVE's survey authoring panel, open the Data Source tab and choose LimeSurvey. Enter your LimeSurvey URL, API username, API password, and the survey ID.
  3. Choose a sync interval (5 minutes, hourly, or manual). SuAVE will poll for new responses and append them to the collection.
  4. Map LimeSurvey question codes to SuAVE column names and annotation qualifiers (e.g., append #number to numeric question codes).
📋

Live classroom use. A common pattern: students fill out a LimeSurvey form during class; the instructor projects SuAVE in the same room. Responses appear within minutes — the class watches the dataset grow and can start filtering before the survey closes.

Column mapping

Both integrations support a column mapping configuration that lets you rename fields and attach SuAVE annotation qualifiers at import time. A mapping file for a typical field survey looks like this:

# kobo_mapping.yaml field_latitude: "Latitude#number#hidden" field_longitude: "Longitude#number#hidden" field_photo: "#img" species_name: "Species" count: "Count#number" observer_notes: "Notes#long" habitat_type: "Habitat"

After mapping, SuAVE automatically enables map view (from the GPS column), image grid (from the photo column), range slider for Count, and long-text search for Notes — all without any manual column editing.

Teaching with SuAVE

SuAVE was designed in parallel with an undergraduate research methods curriculum at UCSD. Several design decisions specifically serve classroom use.

No installation. Students open a URL in any browser — no Python, no R, no data tools to install. Real survey data on day one. Loading the General Social Survey or an archaeological collection takes seconds; students start exploring before the lecture is over. Annotation as homework. The annotation feature lets students document their analytical narrative — filter state plus comment — and submit the URL as an assignment. Live LimeSurvey integration. Students fill out a form; their responses appear as new rows in SuAVE within minutes. The class can watch the dataset grow in real time. Graduated complexity. Beginners use the grid and bars; advanced students connect to Jupyter notebooks. The same interface supports both.

📋

A two-page quick-start handout for students is available on request. Email izaslavsky@ucsd.edu.

REST API

The SuAVE API runs on port 3004 (proxied through the viewer at /api/v1) and provides programmatic read access to survey data. It applies the same filter logic as the browser, so a notebook can receive the ?state= URL parameter from the viewer and query exactly the rows currently on screen.

# List all surveys in the server's surveys directory GET /api/v1/surveys # Survey metadata: facets, item count, enabled views GET /api/v1/surveys/{id} # Filtered items — pass state= from the viewer URL, or individual filter params GET /api/v1/surveys/{id}/items?state={base64state}&limit=5000 # Same endpoint, JSON body for complex filters from Python POST /api/v1/surveys/{id}/items Content-Type: application/json { "filters": { "stringFilters": { "Phylum": ["Mollusca", "Brachiopoda"] }, "numFilters": { "Age#number": { "min": 50, "max": 100 } } }, "limit": 1000, "fields": "Phylum,Age#number,Locality" } # Value distribution for one facet within a filtered subset GET /api/v1/surveys/{id}/facets/{facetName}?state={base64state} # Health check GET /api/v1/health → {"status":"ok","port":3004}

The `:id` segment is the CSV filename without the .csv extension, e.g. suavelocal_BGS_Macrofossils. Items responses include total count, filtered count, pagination offset, and an items array. Format can be JSON (default) or CSV (?format=csv or Accept: text/csv).

The API is part of the suave-api.tar.gz deployment bundle. The full endpoint reference, the Python helper class, and a step-by-step CentOS/RHEL deployment guide (nginx and Apache2 variants) are in DEPLOY.md, distributed with the source on GitHub.

Python client

Python client. suave_api.py (in the suave-api repository) needs only the standard library, with pandas optional; copy it next to your notebook.

from suave_api import SuaveClient c = SuaveClient("https://suave.sdsc.edu/api/v1", "suavedemos_SDG_Indicators_2018", state=None) df = c.items(fields=["Country", "Regional Score (0-100)#number"]) # a pandas DataFrame everything = c.all_items() # pages automatically europe = c.filter_by(**{"Regions for SDGIndex": "Europe"}).items() counts = c.facet_df("Income Group in 2016")

df.attrs carries the total and filtered counts. Passing the state from a launch URL makes the client see exactly what the viewer showed.

Replace a single image

The authoring backend (port 3003) exposes one endpoint for updating an individual item's image without reprocessing the whole collection:

POST /manage/api/surveys/{id}/replace-image Authorization: Bearer {token} Content-Type: multipart/form-data image_id = fossil_042 # current value in the #img column new_image_id = fossil_042_hires # optional — rename in DZC + CSV image = <file> # JPEG/PNG/WebP, max 50 MB → {"ok":true,"image_id":"fossil_042_hires","w":2400,"h":1800,"csv_rows_updated":1}

The endpoint regenerates the full DZI pyramid, patches the DZC manifest in place, and (when new_image_id is supplied) updates every matching value in the survey's #img column. Other items and their tiles are never touched. Authentication requires the same Bearer token or session cookie as the survey management interface — the caller must own the survey.

Self-hosting

The complete SuAVE stack consists of four deployment bundles: the minified front-end (suave2026-dist.tar.gz), the unminified debug front-end (suave2026-dist-debug.tar.gz), the authoring backend (suave2026-backend.tar.gz), and the REST API (suave-api.tar.gz). A fifth bundle (suave-web.tar.gz) contains this website. Deploy order: authoring backend on port 3003, API on port 3004, then serve the static front-end from nginx or Apache2 with proxy rules for those two ports. Full instructions, nginx and Apache2 config templates, PM2 startup, and firewall/SELinux notes are in DEPLOY.md.