Documentation
Everything you need to prepare, publish, and share a SuAVE survey — from CSV formatting to Jupyter notebook integration.
SuAVE (Survey Analysis via Visual Exploration) is a free web platform that turns a CSV file into an interactive, multi-view data explorer for images, maps, charts, and networks. Upload your data, add qualifiers where needed, and SuAVE renders a zoomable grid of image tiles linked to a faceted filter panel — no installation, no code, no database configuration. The platform is hosted at suave.sdsc.edu and is free for academic and non-commercial use.
The open-source SuAVE 2026 frontend (React 18 + Vite) can also be self-hosted and loaded with any CSV directly from the browser or via URL parameters — no account required for local exploration.
#number, #multi, #date, etc. to control how they appearSuAVE works best with 50–50,000 rows. Very small datasets (under 20 items) have limited visual value; very large ones (100,000+) may require server-side pagination — contact us for large-data deployments.
The full reference documentation is at suave-ucsd.github.io/SuAVE-Documentation, covering account setup, image galleries, LimeSurvey integration, bibliographic networks, Jupyter notebooks, and self-hosting. Video tutorials are on the SuAVE YouTube channel.
SuAVE reads standard comma-separated values. The first row is the header — it defines ordinary variable names, reserved variable names, and optional qualifier suffixes (see next section). Values can be text, numbers, dates, pipe-separated lists, or URLs.
For an ordinary variable, the display label is the part before the first #; any following tokens are qualifiers. A column named Geological Period#number#hidden has the label "Geological Period", is treated as a numeric range slider, and is hidden from the filter panel (but shown in the item detail panel). Reserved variable names such as #name are the exception: the entire header is the variable name.
Values in #multi columns are pipe-separated: Cephalopoda|Mollusca|Marine. Each pipe-delimited token becomes an independent checkbox in the filter panel.
UTF-8 encoding is required. Quote fields that contain commas. Empty cells are treated as missing values and displayed as "N/A" in the filter panel. Grouped views (Bucket, Bars, Crosstab and the charts) and the map leave items with no value out by default; tick Display missing values in Settings to show them as an N/A group. On the map they are always drawn, in neutral grey.
Beyond direct upload, SuAVE 2026 accepts data from several sources via URL parameters. A Google Sheets share link (the /edit or /view form) is automatically converted to its CSV export equivalent. Any publicly accessible CSV can be loaded with ?csv_url=URL. A survey config JSON file can bundle the CSV URL, DZC URL, and view settings into a single shareable document (see Survey config files).
SuAVE CSV headers use three distinct conventions: exact reserved variable names, reserved words in spatial variable labels, and qualifiers appended to ordinary variable names. These conventions are not interchangeable.
The following are complete, exact column names. They begin with #, but they are variables—not qualifiers to append to another label.
| Variable name | Purpose | Value in each row |
|---|---|---|
#name | Item display name and detail-panel title | Plain text |
#img | Primary image identifier | Image ID matching the survey's DZI/DZC data |
#href | Makes the #name title clickable | Item-level URL |
#netvis | Associates the item with network-visualization data | Network-data reference used by the Netvis view |
#uniqueid and #views are legacy/internal headers recognized by the original SuAVE loader. Do not add them to new SuAVE 2026 CSV files: row IDs are generated automatically, and enabled views are configured in the survey settings.
Use these exact headers in new CSV files: #name, #img, and #href—not Name#name, Image#img, or Website#href. SuAVE 2026 accepts some labeled forms for compatibility, but the canonical SuAVE convention and the original implementation use exact reserved variable names.
A column tagged #lat, #lon or #geometry is the map's latitude, longitude or shape column, whatever its label says (Широта#lat, Route#geometry#hiddenmore). Without these tags the map looks, as the original SuAVE did, for a label (the part before the first #) containing one of the words below, case-insensitively, so existing surveys keep working. #coord places nothing.
| Or a label containing | Purpose | Recommended header |
|---|---|---|
Latitude | Point latitude | Latitude#number#hidden |
Longitude | Point longitude | Longitude#number#hidden |
Geometry | WKT or GeoJSON geometry for points, lines, or polygons | Geometry#hiddenmore |
Qualifiers are suffixes appended to an ordinary variable name with #. They control the data type, filtering, sorting, and visibility. Multiple compatible qualifiers can be chained, as in Depth#number#hidden.
| Qualifier | Filter panel | Detail panel | Notes |
|---|---|---|---|
(none) | Checkbox list | ✓ | Default for text columns |
#number | Range slider + histogram | ✓ | Values must be numeric |
#date | Date range picker | ✓ | ISO 8601 or MM/DD/YYYY |
#multi | Checkbox list | ✓ | Pipe-separated values; each token is a separate facet |
#long | Search scope only | ✓ | Long text — appears in search scope dropdown, not as checkboxes |
#link | — | Clickable URL | Displayed as a hyperlink in the detail panel |
#hidden | — | ✓ | Hidden from the filter panel and the sort list, visible in detail panel |
#hiddenmore | — | — | Hidden everywhere, the sort list included; useful for internal IDs |
#info | — | Body text | Long description rendered as prose in the detail panel |
#ordinal | Ordered categories | ✓ | Categories whose values have a natural order; numeric-prefixed values are ordered numerically |
#sortquan | Checkbox list | ✓ | Sorts category values by item count instead of alphabetically by default |
#textlocation | — | — | Legacy address-geocoding qualifier; unsupported—use Latitude and Longitude columns instead |
The 2026 parser also accepts #image for #img. #lat, #lon and #geometry mark the map's columns (see Map columns above); #coord is accepted for older files but places nothing.
How images are specified in the #img column depends on which version of SuAVE you are using. The short answer to "can I just paste in a JPEG URL?" is: yes on the hosted platform, no in SuAVE 2026.
The hosted platform supports two image delivery modes. When you upload images through the New Survey dialog, the server processes them into Deep Zoom Image (DZI) pyramids server-side. Your #img column values are then the bare filenames without extension — these are image IDs that the server resolves to the right DZI tile at every zoom level.
For surveys where items already have publicly accessible image URLs (JPEG, PNG), you can also paste those URLs directly into the #img column. The platform will display them at the resolution provided. This works well for moderate-size collections where deep zoom into individual images is not needed.
Images are uploaded through the New Survey dialog — not through a separate tool. When creating or editing a survey at suave.sdsc.edu, the survey settings dialog lets you specify which image files to include. The server processes them into DZI pyramids automatically. There is no longer a separate image upload step via an external service.
Images go up in batches of at most 8 MB, so a slow or briefly dropped connection costs one batch, which is resent automatically, rather than the whole upload. If creating the survey then fails, the images stay on the server and pressing Create again sends only what is missing. File names without their extension become the image IDs, so they must match the #img column; non-Latin names such as Cyrillic are fine. A collection is tiled in one format: JPEG when most images are opaque, with the few transparent ones flattened onto white, PNG when most use transparency. The dialog asks only when the split is unclear.
SuAVE 2026 requires DZI pyramids for all image display. The #img column values must be image IDs, not full URLs. At load time the viewer fetches a DZC manifest file that maps each image ID to its tile location. There is no direct-URL fallback: if no DZC is configured, items without resolved tiles render as colored placeholder shapes.
The DZI format (PNG or JPEG) is read from the DZC XML's Format attribute — SuAVE never hardcodes it, so mixed-format collections work correctly. At high zoom levels in Grid view, SuAVE stitches multiple DZI tiles for seamless full-resolution display. The tile level formula is min(imgMaxLevel, ceil(log2(renderSize))), with no artificial cap.
When you upload images through SuAVE 2026's New Survey dialog, a small composite sprite sheet is generated alongside the DZI pyramid for each survey — a grid of tiny, aspect-preserved thumbnails across a few low zoom levels. The viewer loads this one small file first, so items show a real (if low-resolution) preview immediately, before the full-resolution tiles finish loading — rather than staying blank or tile-by-tile. Surveys migrated in from elsewhere with only a DZI pyramid and no sprite sheet fall back to loading tiles directly; there's no visual difference once tiles finish loading, just a slower initial appearance.
SuAVE detects spatial columns by label (the part before #), not by position. The label must contain the word "latitude", "longitude", or "geometry" (case-insensitive).
Name your columns so the label contains "latitude" and "longitude" — for example Latitude#number#hidden and Longitude#number#hidden. Adding #hidden keeps them out of the filter panel while still enabling the map. Coincident points (identical lat/lon) are automatically clustered at low zoom and spread at high zoom.
A column whose label contains "geometry" is treated as a WKT or GeoJSON geometry column. SuAVE renders polygons, lines, and multipolygon features on the map. The #hidden suffix is recommended to keep the raw WKT out of the filter panel.
A column named country_lat#number#hidden has the label country_lat, which does not contain "latitude" and will not be detected as a spatial column. Spell out the full word.
The survey manager is at /manage. It needs an account: sign in, or use Create account on the sign-in page.
New survey on the dashboard opens a two-step wizard.
Step 1, data source. Choose one:
| Option | What happens |
|---|---|
| Upload CSV file | The file is stored with the survey. The name field is prefilled from the file name. |
| Upload Corpus-DB ZIP | A bibliographic network export; see Corpus-DB surveys. |
| Import CSV from URL once | The server fetches the URL once and keeps its own copy. |
| Link to URL (re-read on every load) | The survey always shows the URL's current contents; the server keeps a copy to fall back on if the URL is unreachable. Not for use with icon collections. |
For Google Sheets, paste the sheet's …/export?format=csv link.
Survey name. The name is shown exactly as typed. The survey's id, which appears in its URL and file names, is derived from it: Cyrillic is transliterated (КОВРЫ22 → KOVRY22), accents are dropped, and other punctuation becomes _. Two of your surveys cannot share an id.
Step 2, configure.
#img value; non-Latin names are fine. Images go up in batches of at most 8 MB, so a dropped connection costs one batch, which is resent automatically. If creating the survey then fails, the images stay on the server: press Create again and only what is missing is sent. Tiles are JPEG when most images are opaque (a few transparent ones are flattened onto white) and PNG when most use transparency; the wizard asks only when the split is unclear. Progress shows uploading, then tile generation, then sprite sheets, and an email is sent when the images are ready. Upload to server chooses among the image servers your installation lists. Save full-size originals for subsequent analysis keeps each uploaded file, unchanged, beside the tiles, at /surveys/<user>_<survey>/originals/<image id>.<ext>, for notebooks and other image analysis; replacing an image in Curate replaces its original too. Without it, full-size images can still be rebuilt from the tiles later (see Curating with an AI assistant).#netvis column.The success screen offers Preview survey and Back to dashboard. A default About page is created, which you can fill in from the editor.
The dashboard shows a card per survey with a thumbnail, title, record count and date.
The editor has seven tabs, each with its own Save, and a Preview link.
US, UK or DRC are recognised. Numeric and date variables are split into 2 to 20 equal ranges.The Curate tab has four parts.
Images lists every picture in the survey's image collection, with its name, id, size in pixels (and megapixels from 1 MP up) and number of zoom levels, and a search box. SuAVE keeps the tile pyramid built from an upload, not the uploaded file, so the size shown is the original's pixel size. The zoom levels are the pyramid's: one per halving, from 1 pixel up to the full size, so ceil(log2(longest side)) + 1; a 1600×1200 picture has 12. Replace… on a picture takes a new file, shows the current and the new picture side by side with their pixel sizes and zoom levels and the new file's size (and a warning if the new one is smaller), and replaces it once confirmed; the result reads, for example, "400×300, 10 levels → 1600×1200 (1.9 MP), 12 levels". The item keeps its place, its id and its data. The new picture is tiled in the collection's own format and its cell in the low-zoom sprite sheet is redrawn, so after a reload of the viewer it shows at every zoom level. If the new file cannot be read, the old picture stays.
Replace several… takes many files at once and pairs each with the picture of the same name, ignoring the extension (IMG_0042.jpg replaces IMG_0042); upper and lower case may differ when that leaves no doubt. Files with no matching picture are listed and skipped: replacement never adds images.
Images and rows. Above the pictures, the tab checks the data's #img column against the collection, as the viewer reads it (the exact id, ignoring spaces around it):
IMG02 for img02, img03.jpg for img03), it is offered with Use it, and Fix all takes every such match. Any other picture can be chosen by its id. Only the #img cells of those rows change; nothing else in the CSV does.N/A, and the names image_not_available and default (the original SuAVE's placeholders) mean "no image" and are only counted.Clones share their source's images, and some older collections serve several surveys. When the collection is shared, the tab names the other surveys and a replacement changes the picture in all of them, after a confirmation. This is allowed only when you own every survey using the collection; otherwise the buttons are disabled. Images hosted on another server cannot be replaced here.
Variables sets how each column is read, as the original SuAVE's Tags dialog did, and rewrites only the header row of the CSV; the data are not touched.
Each column is checked against its values. Under it, a line describes them (how many, how many different, the range of numbers, a few examples), and where the data call for something else a Suggested line says what and why, with Apply. A summary above the table counts the suggestions and offers Apply all suggestions, which takes them all in one click; nothing is saved until Save. For surveys with more than 20 columns, a search box finds variables by name and Only columns with suggestions, problems or changes hides the rest (the General Social Survey's 796 columns come down to about 30). The table shows 50 variables at a time, with Previous and Next above and below it; search and the filter apply to all of them first, and Apply all suggestions and Save always cover every column, not just the page shown. The suggestions are:
|, and Long text for long, mostly different passages (a few long category labels stay categories).#name) is never suggested Hidden or By count.lat, lng and similar labels to Latitude and Longitude when their numbers are coordinates, so the map finds them.Qualifiers are checked against what the viewer actually reads. A Number or Ordinal value is read from its leading number, so "89 or older" counts as 89 and "2nd important" as 2; only a value with no leading number, such as "$75+", is missing. A Date value is read if it starts with a four-digit year or the browser reads it, as with "11/21/24" or "Nov 21, 2024". An Ordinal of up to 12 ordered words ("agree", "disagree"…) is not flagged: the filter panel keeps their order, though Parallel coordinates can only plot numbers. When at least 80% of the values can be read, the qualifier is kept with a note saying how many cannot; below that it is flagged in red and another type is suggested. Types you chose that the data allow, such as Description or Item URL, are left alone.
#sortquan): values listed by how many items have them.#hidden: not in the filter panel or the sort list, still in the info panel) or Hidden everywhere (#hiddenmore: not in the info panel either).#img, #name, #netvis), and the structural columns that stand without a label, such as #href (the page opened from an item's title) or #info (its description). As in the original SuAVE, a header that starts with # is not a variable, so such columns need no label. Change them by replacing the CSV. Qualifiers SuAVE does not offer here, such as #title, are kept as they are.#number, #date or #ordinal names nothing. Creating a survey or replacing its CSV with such a header is refused, with the columns listed (the original SuAVE ignored such columns); one already in a survey is hidden in the viewer and shown here as needs a label, so you can name it, for example Price#number.#lat), Longitude (#lon) and Map shape (#geometry), whatever their labels; otherwise, as in the original SuAVE, it looks for "latitude", "longitude" and "geometry" in the labels. The review suggests the marks for coordinate or shape columns the labels would not reveal (such as lat, lng or Route), and Hidden everywhere for long shapes. #textlocation and #coord place nothing; they are kept where a file has them but not offered.| Choice | Qualifier | Effect |
|---|---|---|
| Text | none | a list of values to tick in the filter panel |
| Number | #number | histogram and range slider |
| Date | #date | filtered by date range |
| Long text | #long | searchable from the toolbar search box, not listed in the filter panel |
| Link | #link | a link in the info panel |
| Item URL | #href | the page opened when the item title is clicked |
| Ordinal | #ordinal | whole numbers 0 to 9: a list of values in the filter panel, a numeric axis in Parallel coordinates |
| Multiple values | #multi | several values in one cell, separated by a vertical bar |
| Latitude, Longitude | #lat, #lon | the map's coordinates, whatever the label |
| Map shape | #geometry | WKT or GeoJSON lines and areas for the map, whatever the label |
| Description | #info | the item's description in the info panel |
Values looks at what the cells hold and rewrites only the cells you change; the header and every other record stay as they were, byte for byte. Columns that are not filters (images, names, links, descriptions, long text, map shapes and coordinates as text) are left out. Each column lists what it found, with a suggestion where one is clear:
-, NA, n/a, null or ?: the viewer shows them as a value of their own. Suggested: make them missing. Words that may be real answers ("None", "Don't know", "Refused") are suggested one at a time, never in Apply all. In columns of country or state codes, NA and DK are left alone.A summary counts the suggestions and Apply all takes the unambiguous ones. Any value can also be changed with Replace… or Make missing, and Show all N values lists every value of a column with its count. Columns are paged 50 at a time with a search box and Only columns with findings or changes. Nothing is written until Save. As with Variables, a survey with Keep link must be imported first.
History lists every version of the survey's data, newest first: when, by whom, rows, columns and size, and what changed in plain words, for example "Values: “-” → missing in Kingdom (4 cells)" or "Replaced the data with the file “birds.csv” · 60 → 58 rows". Every change to the CSV is recorded, whether made in Curate (Variables, Values, image references) or by replacing the CSV in the Info tab. The data a survey had before its first recorded change are kept as version 1.
prov:Entity, each change a prov:Activity that used the previous version and generated the next, by a prov:Person. The sentences in the list are written from this record.An AI assistant such as Claude Code or OpenAI's Codex can do the Curate tab's work in conversation: review a survey's variables, values, images and history, explain what it finds, and make the changes you approve. It works through SuAVE's MCP server (the Model Context Protocol, the standard way assistants use outside tools) with an access token instead of your password.
In the gallery, AI assistant opens the setup. The simplest way is to connect by address, https://suave.sdsc.edu/manage/api/mcp, with nothing to install:
claude mcp add --transport http suave -s user https://suave.sdsc.edu/manage/api/mcp, then /mcp in Claude Code, choose suave, Authenticate.codex mcp add suave --url https://suave.sdsc.edu/manage/api/mcp, then codex mcp login suave.The sign-in page names the app and the site it will return to; continue only if you started the connection there. The connection appears in the panel's list ("Claude (connected app)") and can be revoked there. It renews itself for 90 days of use.
The SuAVE MCP server can also run on your own computer, with an access token instead of the sign-in (the panel's "Or run SuAVE's MCP server on your computer"):
suave-mcp folder in your home folder: the panel gives the terminal command, mkdir -p ~/suave-mcp && mv ~/Downloads/suave-mcp.js ~/suave-mcp/.claude mcp add suave … for Claude Code or codex mcp add suave … for Codex; it already contains the token. For Claude Desktop, put the panel's settings into Settings → Developer → Edit Config, with the two paths filled in. If SuAVE was added before, remove it first (claude mcp remove suave -s user, codex mcp remove suave).Then start the assistant again and ask, for example, "List my SuAVE surveys" or "Review the values in my survey Birds and suggest fixes". The assistant can list your surveys, read each one's variables (with a profile of every column), problems in the values, images against rows, and versions, and compare a version with the current data. It changes qualifiers and labels, values, image references, or restores a version, only with an explicit list of changes.
It can also look at the pictures. get_survey_definition says what a survey is made of: rows and columns, whether its images are on this server, their format and sizes (smallest, median and largest side, how many over 2000 px), zoom levels, and how many full-size images are kept. get_image shows it one picture, whole and scaled down or a region at full resolution (up to 2048 px across), so it can examine detail such as brushwork or handwriting. reconstruct_full_images_from_dzi writes the full-size image of chosen images, or all, rebuilt from their zoom tiles (the top zoom level has each image's full pixel size), into the same originals/ folder; it runs in the background, never overwrites an uploaded original, and list_full_images gives the files' addresses for a notebook. Larger analyses, such as comparing brushstrokes across thousands of paintings, belong in a notebook that reads those files.
A connected app or a token can only curate your surveys: it cannot delete them, change their settings, upload files, sign in or make other tokens, and administrator rights never pass through it. Every change it makes is saved in the survey's History as made by that token on your behalf ("you via “Claude Code on my laptop”"), so any change can be compared and restored. Tokens expire after a year; Revoke… stops one at once. Never paste a token into a chat or a shared document; if one was, revoke it and make another.
| You have | Choose | Notes |
|---|---|---|
| Photos, scans or artwork | Upload images | #img holds each image's file name without extension |
| A tile collection made elsewhere | Link to existing DZC | Paste the .dzc URL |
| No images, but categories | Use icon collections | Then set Shape/Color variables in Icons |
| Countries | Use icon collections → Country flags | Values may be names or ISO codes |
| A Corpus-DB export | Upload Corpus-DB ZIP | Flags are set automatically from the authors' countries |
| Images arriving from a form | Create empty repository | See KoboToolbox and LimeSurvey |
You can copy any survey you can open into your own account: public surveys, your own, and private surveys that list your email. Use Clone to my gallery in the viewer's toolbar (or in the Jupyter dialog). After signing in you are asked for a name, "<name> (clone)" by default, up to 200 characters, and the copy opens in the editor.
/gallery/<username> listing its visible surveys, with Open and About buttons.SuAVE 2026 can open any CSV directly from the landing page: drag and drop a file, or paste a URL (Google Sheets links are converted automatically). No account is needed, and a dropped file never leaves your browser. Icon-based and text-only surveys work fully this way; image collections need a configured tile (DZC) URL.
Corpus-DB is a free service that builds global bibliographic networks from OpenAlex: you describe a network by a thematic publication scope or by a list of authors, refine it by keywords, countries, institutions and publication years, and download it as a dataset. Imported into SuAVE, it becomes a survey of authors that you can filter, map, compare, and explore as a co-authorship network, following how it grows over time.
The whole workflow has four stages: generate, clean, (optionally) merge, and publish.
Go to corpus-db.sdsc.edu and sign in with your SuAVE username and password. Choose Search, check that the email at the top of the page is yours (results are sent there), and give the project a name. Then choose one of two search types.
Scope (important terms). Use this to build a network around a topic.
UC San Diego, UCSD, University of California San Diego) so that nothing relevant is missed.Author OpenAlex ID. Use this when you already know the people: enter their OpenAlex author ids, comma-separated. Options are the same as above (keywords, years, institutions, co-author threshold), plus whether to include external authors in general, or only those within the authors' institutions.
Press Submit. Depending on the size of the search, an email arrives within seconds, minutes or, for very large ones, hours. Follow its link to corpus-db.sdsc.edu/collect, enter the code from the email, and the dataset downloads. If the first submit shows an error, open a new tab, return to Corpus-DB and sign in again.
The download contains three files: the authors CSV, a *-publications.csv, and the network JSON. If they arrive as a folder rather than a ZIP, put all three into one ZIP for SuAVE.
Bibliographic data repeats people under slightly different names. Clean it before publishing:
For harder cases, the cleaned authors CSV can be curated further in OpenRefine: load it, cluster the Name column (Edit cells → Cluster and edit), merge clusters so that one person has one spelling (first name then last name, spelled out where possible), try other clustering methods including an n-gram size of 1, remove rows such as "Anonymous" or "Author" with a text facet, and finally add a last-name column (value.split(" ")[-1]) and review clusters by last name, merging only where the other fields agree. Save as UTF-8.
The CorpusDB Merge tool combines several datasets or networks. Its operations are: import CSV, merge CSVs (on a key such as the OpenAlex id, choosing whether overlapping rows are overwritten), conjoin networks, add a facet (column) to a CSV, edit or delete facets, build a network from chosen facets, and export the result as CSV, JSON or a ZIP for SuAVE. Steps are connected visually, by dragging from one block's tab to the next; save the network before leaving. A typical use is to conjoin two networks over the same people, for example co-authorship and shared research concepts, so that each author node links both to collaborators and to topics. The tool's Help button has a video tutorial and worked examples.
In the survey manager choose New survey → Upload Corpus-DB ZIP, select the ZIP, name the survey and create it. What the import does:
.csv in the ZIP that is not *-publications.csv) becomes the survey. Each row is one author, with columns such as the OpenAlex id, affiliation, city, region and country (with coordinates), OpenAlex concepts, total publications, first and latest publication dates, citedness, H index, the search scope and keywords, collaborators in scope, publication dates, a Show Publications cell, #img (the author's country code) and #netvis (the author's key in the network).#img holds ISO country codes (un where unknown), matched to SuAVE's shared flag collection.*-publications.csv), the number of authors, the span of publication years, the search scope, its keywords and the most prominent OpenAlex concepts. Replace it any time from the editor's About tab.Show Publications. Each author's Show Publications button, in the info panel, opens a page listing that author's publications within the scope, retrieved live from Corpus-DB.
Author photos instead of flags. A notebook in the SuAVE documentation finds candidate photos of authors from their name, affiliation, city and country. To use photos, set each author's #img value to their photo's file name (without extension), and either create the survey with Upload images in step 2 of the wizard, selecting the photos, or attach an existing tile collection from the editor's Info tab.
Exploring. Map the authors and colour them by H index or citedness; compare countries and concepts in Crosstab or Heatmap; follow first and latest publication dates in Timeline; and in Netvis, move the year range to watch the network grow, trace shortest paths between two authors, and open node statistics to find brokers and central figures.
Step-by-step guides with screenshots are on the SuAVE documentation site: generating a dataset, cleaning it, curating names in OpenRefine, the Merge tool, conjoining networks and author photos.
The SuAVE team maintains a curated set of sample datasets for workshops and experimentation, covering three broad types.
These use icon specifications (shape and color driven by variable values) rather than photographs. The EarthCube Member Survey and various NSF award databases fall into this category. They are the simplest starting point: prepare a CSV, annotate the columns, upload — done.
Surveys like the Observing Systems Explorer and 2018 SDG Indicators dataset use small, bounded image sets that are easy to process. Good for learning the image upload workflow without managing thousands of files.
Examples include the Picasso Paintings survey, the USGS Earth As Art collection, the Van Gogh paintings collection, BGS Macrofossils, and wetland soil samples. These demonstrate the full DZI tile pipeline and deep-zoom interaction.
Each dataset is a public Google Drive folder with a README describing the data and its source, the CSV, and, where the survey has them, the images, ready to publish in SuAVE. The same list is kept in the SuAVE sample datasets document.
| Dataset | Shows how to publish |
|---|---|
| Picasso paintings | High-resolution images |
| Earth as Art (USGS) | High-resolution images |
| EarthCube member survey 2013 | No images: icons from variables |
| Earth Observing Systems | A small set of icons and logos |
| Wetland samples, Lower Mekong | High-resolution images |
| San Diego vacant lots | High-resolution images |
| 2018 SDG Indicators | A small set of icons, and a map |
| Jordan food | Images |
Published versions of many demo surveys can be browsed in the 100+ Surveys gallery.
A good first exercise: go to the NSF Awards search, run any keyword search, export the results as CSV, add a #number qualifier to the Amount column and a #date qualifier to the start date, then upload. You get a funded-project browser with a histogram and date filter in under ten minutes, no images required.
Every view shows the same filtered set of items; switching views keeps the filters (and closes the info panel). How a view treats items with no value for its variable is controlled by Settings → Display missing values (off by default): when off, such items are left out of grouped views; when on, they form an N/A group, placed last.
Every filtered item as a square tile, in the toolbar's sort order. At zoom 1 the whole set fits the canvas; zooming keeps the column count and enlarges the tiles, up to the resolution of the source images.
A histogram of image stacks: one column per group, tiles stacked from the bottom, column height proportional to the count.
Albania–Chile; if every value is a number, numeric ranges instead.One vertical column of tiles per group, side by side. Each column scrolls on its own and the row of columns scrolls sideways, so every group can be browsed in full.
A two-way table: one variable down the rows, another across the columns. Each cell shows thumbnails of its items under a count overlay coloured by size.
Items as points, lines or areas on an OpenStreetMap base map.
#) contains "latitude" and one containing "longitude" give points; a label containing "geometry" gives WKT lines and polygons. #hidden columns work, e.g. Latitude#number#hidden. Without such columns the view says so.N/A legend entry, with a count, appears when Display missing values is on.A spreadsheet of the filtered items: a thumbnail column (when the survey has images), then every visible column in CSV order. Hidden columns and coordinates are left out. Cells show up to five lines and scroll inside.
A compact list in the toolbar's sort order. Each row shows a thumbnail, the name and one line of description: the first #info column, otherwise the first #long column. Click a row to open the info panel.
A graph built on the fly from the survey's own columns.
|.A pre-computed network supplied with the survey as JSON (typically a co-authorship network from a Corpus-DB export). The tab appears whenever the survey has a #netvis column.
#netvis value is a key into the JSON's data. Each entry lists the item's connections (each with a name, category, optional weight, year list, icon and links). The item itself is the node named by its #name value. The graph is rebuilt from the filtered items, so filters shrink it.json { "config": { "version": 2, "view": { "<network>": { "slider_label": "Year", "node_scaling_options": ["H Index"], "max_weight": 10, "connection_title": "Co-authors of %s" } } }, "data": { "<key from the #netvis column>": [ { "category": "Author", "network": "<network>", "slider_variable": [2019, 2021], "icon": "assets/flags/US.png", "links": ["https://openalex.org/…"], "connections": [ { "name": "Jane Doe", "category": "Author", "weight": 3, "slider_variable": [2020], "icon": "assets/flags/FR.png" } ] } ] } } ` Files without config.version` are read as the older v1 layout (one entry per key).A colour matrix of two variables.
Two numeric variables against each other.
Item counts over time, as stacked bars.
One line per item across parallel numeric axes.
Flows between categories.
Click an item in any view to open it.
#name value, a link when the survey has an #href column.#info column, as formatted text (HTML is cleaned before display).#link values are links. A cell containing getPublication({...}) becomes a Show Publications button (see Corpus-DB surveys).On phones these move to a bar at the bottom.
A multi-condition filter. Match All (AND) or Any (OR). Text conditions: contains, exactly equals, does not contain; numbers: between; dates: from/to. In All mode the conditions become ordinary filters with chips; in Any mode the matching set is applied directly, without chips, and is not part of share links.
Builds a link that reproduces the current view, sort, filters, search and selected item, with a Copy button and shortcuts for Facebook, LinkedIn, WhatsApp, X and email. Calculated columns are not included.
Shown for surveys stored on the server. Opens the survey manager with a copy of this survey ready to be made (see Cloning a survey).
Saves the filtered items as {survey}_filtered.csv, with the original columns in order, dates as YYYY-MM-DD and missing values as empty cells.
Builds a new numeric column from existing ones. Each row is a number column or a constant, optionally transformed (Σ sum over all rows, −x, √x); rows are joined by + − × ÷ and evaluated left to right, without operator precedence. A missing input, division by zero or the square root of a negative gives a missing result. The column exists for this session only.
Saves a comment with a snapshot of the current filters and search as a named pattern in this browser. Saved patterns can be restored (filters and search are re-applied) or deleted. Unlike per-item notes, a pattern describes a subset.
Changes apply when you press Apply; all but the survey name are remembered in this browser.
| Setting | What it does |
|---|---|
| Survey name | Display name for this session only |
| Display missing values | Show items with no value as an N/A group in grouped views and the map legend; the hint counts them for the current variable |
| Tile labels | Always, Hover or Never, and the Label column |
| Default sort order | Ascending or descending |
| Info panel | Shrink canvas or Overlay |
| Visible variables in filter panel | Move variables between shown and hidden lists |
| Lightbox fields | Up to 8 fields shown in the Grid lightbox |
Shows the survey's About page in a dialog. Escape closes it.
For Heatmap, Scatter, Timeline, Parallel and Sankey: saves the chart at twice screen resolution.
Fills the screen; Escape leaves.
Add ?lang= to the viewer URL to choose a language: ar_AR, es_MX, ka_GE, kk_KZ, ru_RU, tr_TR, zh_CN or zh_TW. English is the default, and anything not yet translated stays in English.
The filter panel lists every facet column as a collapsible accordion. Checking or unchecking values updates all views in real time — no page reload. Multiple selections within one facet combine as OR; selections across different facets combine as AND.
The toolbar search box searches across all columns by default. When a #long column is selected in the scope dropdown, search narrows to that field. Within each facet accordion, a separate "Filter values…" input filters the checkbox list display — it does not affect item counts.
The search icon in the toolbar opens the multi-condition filter builder. Match all conditions (AND) or any (OR). Text conditions are contains, exactly equals and does not contain; numbers take a range, dates a from/to. See Info panel, toolbar & settings.
A row of colored chips below the toolbar summarizes all active filters. Click × on any chip to remove that filter. A Clear All button in the filter panel header removes everything at once.
The annotation icon saves the current filter state — every active checkbox, range, and search term — as a named pattern in the browser's localStorage. Restore any saved pattern with a single click. Useful for teaching, where students document their analytical steps.
The share icon builds a link that reproduces the current view, sort order, filters, search and selected item, with a Copy button and shortcuts for Facebook, LinkedIn, WhatsApp, X and email. Calculated columns are not included.
The download icon exports the currently visible items as a CSV file — only rows that pass all active filters are included. Reserved variable names and qualifiers are preserved in the header.
For chart views (Heatmap, Scatter, Timeline, Parallel, Sankey), a camera icon in the toolbar downloads the current chart as a high-resolution PNG.
A survey config file is a plain JSON document that bundles all the information needed to open a survey without a backend account. It is useful for self-hosted deployments, for sharing a survey with a specific pre-selected view or overlay, or for attaching Streamlit apps to a CSV that lives on an external URL.
Only csv is required. All other fields are optional. Load a config file with the URL parameter ?surveyconfig=URL. Google Sheets share links in the csv field are automatically converted to their export equivalent.
SuAVE 2026 accepts several URL parameters for opening a specific survey directly.
| Parameter | Purpose | Example |
|---|---|---|
?survey=filename.csv | Load a CSV from the server's surveys directory | ?survey=suavelocal_BGS_Macrofossils.csv |
?user=X&file=Y | Original SuAVE URL style; equivalent to ?survey=X_Y.csv | ?user=suavelocal&file=BGS_Macrofossils |
?csv_url=URL | Load any CSV from an arbitrary public URL | ?csv_url=https://…/data.csv |
?surveyconfig=URL | Load a JSON config file (CSV + DZC + views + metadata) | ?surveyconfig=/configs/my.json |
?lang=CODE | Open the interface in another language (see Interface languages) | ?lang=ru_RU |
?skip_preamble=true | Suppress the intro modal on surveys that have one |
Add ?lang=CODE to any survey URL and SuAVE opens in that language. Only the interface is affected: labels, buttons, menus, dropdowns, and the inline help lines under each view. Survey data is never translated, so variable names and values appear exactly as they are in the CSV.
Currently available:
| Language | Code | Native name |
|---|---|---|
| English (default) | en_US | English |
| Arabic | ar_AR | العربية |
| Chinese (Simplified) | zh_CN | 简体中文 |
| Chinese (Traditional) | zh_TW | 繁體中文 |
| Georgian | ka_GE | ქართული |
| Kazakh | kk_KZ | Қазақша |
| Russian | ru_RU | Русский |
| Spanish | es_MX | Español |
| Turkish | tr_TR | Türkçe |
A complete example:
The parameter combines with every other URL parameter, and it survives the share link, so a link you copy while working in Kazakh reopens in Kazakh for whoever you send it to. Leaving the parameter off gives you English, and so does an unrecognized code.
The list grows as people ask for it. If you would like to work with SuAVE in a language that is not here, write to izaslavsky@ucsd.edu and we will send you the file to fill in. A translation is one JSON file of short interface strings, with the English on the left and your language on the right, and it takes an afternoon rather than a project.
Self-hosted deployments can add one directly. Drop the finished file into the frontend's public/ directory under its language code, say fr_FR.json, and it is served by name the first time someone passes ?lang=fr_FR. Nothing needs to be registered and the application does not need rebuilding.
SuAVE hands the survey you are looking at, filters included, to Python: to a Jupyter notebook for analysis, to a Streamlit app, or to any script through the REST API. Results can come back as a new survey.
The Jupyter and Streamlit buttons appear beside the view tabs when two things are true: the survey has jupyter or streamlit ticked in the editor's Views tab, and the server's defaults.json lists at least one Jupyter or Streamlit server. The server list is set by the administrator for the whole installation; authors only switch the buttons on or off per survey. Private surveys are not available to these integrations.
The Jupyter dialog has up to three tabs.
izaslavsky/suave-notebooks, main, SuAVEDispatch.ipynb; your choice is remembered). SuAVE stores the current survey and filters for 30 minutes under a one-time token, starts the repository on mybinder.org, and the notebook picks the parameters up automatically. The repository must include SuAVE's small receiver (suave/receiver.py), as suave-notebooks does.SUAVE_TOKEN and SUAVE_HOST, with Copy buttons, to paste into the notebook's first cell. The token also lasts 30 minutes.The dialog also offers Clone to my gallery.
The suave-notebooks collection. SuAVEDispatch.ipynb is a menu of ready-made analyses that run on the survey you sent:
| Group | Notebooks |
|---|---|
| Statistics | descriptive statistics, contingency tables, factor contributions, supervised classification and regression, PCA, clustering, outlier detection |
| Arithmetic and wrangling | derived variables, variable transforms |
| Spatial | geographically weighted regression, exploratory spatial analysis, aggregate maps |
| Networks | network building and metrics for Netvis surveys |
| Images | colour statistics, image classification; AI captioning (BLIP), object detection (DETR), image clustering (CLIP) |
| Text | sentiment, zero-shot classification, geocoding, topic modelling, entity linking to Wikidata, OpenAlex enrichment |
Image notebooks need the survey's full-size images: uploaded originals, or full-size images rebuilt from the tiles (see Curating with an AI assistant), at /surveys/<collection>/originals/; network notebooks need a #netvis column.
Publishing back. A notebook can create a new SuAVE survey from its results. It asks for your SuAVE username and password and creates the survey in your account.
The Streamlit dialog lists the installation's Streamlit servers; Connect opens the chosen app with the survey (user, CSV, image collection, survey URL, API address and filter state) in its URL. The standard launcher, suave-launcher, forwards to:
The arithmetic and GWR apps can publish their result as a new survey, after you sign in with your SuAVE account.
SuAVE connects directly to KoboToolbox and LimeSurvey through their REST APIs. Once configured, new survey submissions flow into SuAVE automatically — no CSV exports, no scheduled imports, no reformatting.
KoboToolbox is the standard tool for humanitarian and field research surveys. It supports GPS location capture, photo uploads, offline data collection, and branching logic — making it the right choice for ecology surveys, archaeological fieldwork, urban assessments, and any study where respondents are not at a desk.
The integration uses KoboToolbox's REST service feature. To connect a Kobo form to SuAVE:
https://suave.sdsc.edu/kobo/submit?user=YOUR_USERNAME&file=YOUR_SURVEY.csvGPS and images map automatically. Kobo geopoint fields are split into columns whose names contain Latitude and Longitude, so map view activates without additional qualifiers. Photo fields populate the reserved #img variable, which supplies the image grid.
LimeSurvey is widely used for course-integrated research and institutional questionnaires. The SuAVE–LimeSurvey bridge uses LimeSurvey's RemoteControl API to pull responses on a configurable schedule (or on demand from the SuAVE authoring panel).
#number to numeric question codes).Live classroom use. A common pattern: students fill out a LimeSurvey form during class; the instructor projects SuAVE in the same room. Responses appear within minutes — the class watches the dataset grow and can start filtering before the survey closes.
Both integrations support a column mapping configuration that lets you rename fields and attach SuAVE annotation qualifiers at import time. A mapping file for a typical field survey looks like this:
After mapping, SuAVE automatically enables map view (from the GPS column), image grid (from the photo column), range slider for Count, and long-text search for Notes — all without any manual column editing.
SuAVE was designed in parallel with an undergraduate research methods curriculum at UCSD. Several design decisions specifically serve classroom use.
No installation. Students open a URL in any browser — no Python, no R, no data tools to install. Real survey data on day one. Loading the General Social Survey or an archaeological collection takes seconds; students start exploring before the lecture is over. Annotation as homework. The annotation feature lets students document their analytical narrative — filter state plus comment — and submit the URL as an assignment. Live LimeSurvey integration. Students fill out a form; their responses appear as new rows in SuAVE within minutes. The class can watch the dataset grow in real time. Graduated complexity. Beginners use the grid and bars; advanced students connect to Jupyter notebooks. The same interface supports both.
A two-page quick-start handout for students is available on request. Email izaslavsky@ucsd.edu.
The SuAVE API runs on port 3004 (proxied through the viewer at /api/v1) and provides programmatic read access to survey data. It applies the same filter logic as the browser, so a notebook can receive the ?state= URL parameter from the viewer and query exactly the rows currently on screen.
The `:id` segment is the CSV filename without the .csv extension, e.g. suavelocal_BGS_Macrofossils. Items responses include total count, filtered count, pagination offset, and an items array. Format can be JSON (default) or CSV (?format=csv or Accept: text/csv).
The API is part of the suave-api.tar.gz deployment bundle. The full endpoint reference, the Python helper class, and a step-by-step CentOS/RHEL deployment guide (nginx and Apache2 variants) are in DEPLOY.md, distributed with the source on GitHub.
Python client. suave_api.py (in the suave-api repository) needs only the standard library, with pandas optional; copy it next to your notebook.
df.attrs carries the total and filtered counts. Passing the state from a launch URL makes the client see exactly what the viewer showed.
The authoring backend (port 3003) exposes one endpoint for updating an individual item's image without reprocessing the whole collection:
The endpoint regenerates the full DZI pyramid, patches the DZC manifest in place, and (when new_image_id is supplied) updates every matching value in the survey's #img column. Other items and their tiles are never touched. Authentication requires the same Bearer token or session cookie as the survey management interface — the caller must own the survey.
The complete SuAVE stack consists of four deployment bundles: the minified front-end (suave2026-dist.tar.gz), the unminified debug front-end (suave2026-dist-debug.tar.gz), the authoring backend (suave2026-backend.tar.gz), and the REST API (suave-api.tar.gz). A fifth bundle (suave-web.tar.gz) contains this website. Deploy order: authoring backend on port 3003, API on port 3004, then serve the static front-end from nginx or Apache2 with proxy rules for those two ports. Full instructions, nginx and Apache2 config templates, PM2 startup, and firewall/SELinux notes are in DEPLOY.md.