# KBase Documentation

Here you will find guidance for getting started with KBase, full documentation for users and App developers, and help with troubleshooting.

![](/files/-Lv7HXxdyn1RB98wBTNC)

## How to cite KBase:

Wood-Charlson EM, Henry CS, Dehal PS, Mahmud G, Allen BH, Beilsmith K, et al. KBase: Open-source Platform for Collaborative Biological Data Analysis and Publication. Journal of Molecular Biology. 2026; 169676. doi:[10.1016/j.jmb.2026.169676](https://doi.org/10.1016/j.jmb.2026.169676)

Arkin AP, Cottingham RW, Henry CS, Harris NL, Stevens RL, Maslov S, et al. KBase: The United States Department of Energy Systems Biology Knowledgebase. Nature Biotechnology. 2018;36: 566. doi: [10.1038/nbt.4163](https://www.nature.com/articles/nbt.4163)

## Funding:&#x20;

This work is supported as part of the Genomic Sciences Program DOE Systems Biology Knowledgebase (KBase) funded by the [U.S. Department of Energy, Office of Science, Office of Biological and Environmental Research ](http://science.energy.gov/ber/)under Award Numbers DE-AC02-05CH11231, DE-AC02-06CH11357, DE-AC05-00OR22725, and DE-SC0012704.

<div align="center"><img src="/files/-MFXpJoqUzxUiUY_DNHe" alt=""></div>


# KBase Terms & Conditions

For links to policies and information on KBase resource availability.

KBase is free to use and funded by the U.S. Department of Energy. When signing up for a KBase account, you accept the [KBase Terms & Conditions](https://www.kbase.us/about/terms-and-conditions-v2/), which includes our Use Agreement, Code of Conduct, Policies, and Service Level Agreement.&#x20;

### Use Agreement & Code of Conduct

Users will follow the [KBase Use Agreement](https://www.kbase.us/about/terms-and-conditions-v2/#use_agreement/), which outlines the above policies around data and privacy and the [KBase Code of Conduct](https://www.kbase.us/kbase-code-of-conduct/), which outlines expectations of how users and the KBase team respect each other and collaborate.  &#x20;

### KBase Data Policy

KBase follows the [Information and Data Sharing Policy](http://genomicscience.energy.gov/datasharing/) of the Genomic Science Program of the Office of Biological and Environmental Research within the Office of Science. See the [KBase Data Policy & Sources](https://www.kbase.us/about/terms-and-conditions-v2/#data_policy) page for details, including information on publicly available data.&#x20;

### Privacy Policy

*All data uploaded by users is private to them unless they choose to share it*. KBase does not release or use private data for any internal analyses, with the exception of collecting metrics on 1) amount of data and 2) the distribution of data types in system, and 3) app usage across system as an aggregate of user activity. See the [KBase Privacy Policy](https://www.kbase.us/about/terms-and-conditions-v2/#privacy_policy) for details. &#x20;

### Service Level Agreement & Compute Resources

KBase compute resources are available in a [fair-sharing system](https://batchdocs.web.cern.ch/fairshare/fairshare.html#:~:text=Fair%2Dshare,belong%20to%20the%20same%20group) that enables equitable access for all users to KBase compute resources. Submitted jobs, will be staggered in the job queue with other users based on [user priority](https://htcondor.readthedocs.io/en/latest/users-manual/priorities-and-preemption.html). This means that some jobs may take longer to finish based on previously submitted jobs and available machine slots.&#x20;

To check the status of the job queue and wait times for KBase Apps, view [the Job Log](/getting-started/narrative/job-browser#job-log).&#x20;

### Licensing

**Software**  ⏤ Software developed by the KBase project team or contributed by community developers to KBase, is stored and maintained in the public KBase GitHub code repository under the MIT Open Source License (“License”).


# Getting Started

An introduction to KBase and guide on how to get started.

What you can find in this section:

1. [Signing Up and Signing In](/getting-started/sign-up)
2. [KBase Supported Browsers](/getting-started/browsers)
3. [A Quick Start Guide to KBase](/getting-started/quick-start)
4. [KBase Narrative Interface User Guide](/getting-started/narrative)
5. [FAQs](/getting-started/faq)


# Signing Up and Signing In

Sign up and get started using KBase

## **Signing Up**

Create a free KBase account using one of our supported identity providers: Google, [ORCiD <img src="/files/-M9K9Aepzk1FNZ3jYzOv" alt="" data-size="line">](https://orcid.org/) , or [Globus](https://www.globusid.org/login). For detailed guidance, see the [Step-by-Step Signup Guide](/getting-started/sign-up/step-by-step).

After you have established a sign-in through your Google, ORCiD, or Globus account, you will be able to  use the KBase Narrative Interface to perform your computational systems biology analyses.

{% hint style="info" %}
KBase is an ORCiD [Member Organization](https://orcid.org/members/0010f00002IL95pAAD-lawrence-berkeley-national-laboratory). Sign in or link your[<img src="/files/-M9K9Aepzk1FNZ3jYzOv" alt="" data-size="line">to KBase](/manage-account/link-orcid).&#x20;
{% endhint %}

### **Terms and Conditions**

KBase is sponsored by the U.S. Department of Energy and is free for all to use (even if you are not in the United States). When you link an identity to KBase, you will need to accept our [Terms & Conditions](https://www.kbase.us/about/terms-and-conditions-v2/), including our [User Agreement](https://www.kbase.us/about/terms-and-conditions-v2/#use_agreement). You may also want to review our [Privacy Policy](https://www.kbase.us/about/terms-and-conditions-v2/#privacy_policy) and [Data Policy](https://www.kbase.us/about/terms-and-conditions-v2/#data_policy). Users shall follow the [KBase Code of Conduct](https://www.kbase.us/kbase-code-of-conduct/).&#x20;

[Sign up for a free KBase account](https://narrative.kbase.us/#signup)

## **Signing In**

Click the Sign In button from the KBase website or KBase Narrative Interface. Select the identity provider you chose when you signed up for your KBase account, and use the credentials as you would  when signing into the that provider's website. You can [link multiple accounts to sign into KBase](/manage-account/link-accounts).&#x20;

## How do I set or reset my password?

If you signed up for KBase and your password does not seem to work, you may reset your password at any time with the identity provider who manages it. KBase itself does not manage passwords of its own, so all password management must be done through your identity provider.

If the last time you logged into your KBase account was prior to June 2017, see the [authentication update page](/getting-started/sign-up/auth-update).

{% hint style="danger" %}
If your authentication provider is managed by someone else, you may lose access to your KBase account. It is strongly recommended to attach a secondary authentication method you control to your account.
{% endhint %}

####


# Step-by-Step Sign Up

This page provides step-by-step instructions for signing up to use KBase. It is recommended to use your [ORCID](https://orcid.org/) identifier to create a KBase account.  However, you can also use your existing Google or [Globus](https://www.globusid.org/login) account and it will be linked to your new KBase account.

1\. From the kbase.us home page, click the [Sign Up](https://narrative.kbase.us/#signup) button in the top right corner.

<img src="/files/egjgxiLRuEctIcIOJ6ht" alt="" width="563">

2\. Choose your existing [ORCID](https://orcid.org/), Google, or [Globus](https://www.globusid.org/login) account to sign up for a KBase account. We recommend signing up with your ORCID identifier to all of the publishing features of KBase.

<figure><img src="/files/lB9SjPIqZivaiYJhAB9Z" alt="" width="563"><figcaption></figcaption></figure>

{% hint style="info" %}
**What if I don't have a ORCiD, Google, or Globus account?**

It’s easy and free to get a new account with one of our identity providers. <br>

[Get an ORCID identifier](https://orcid.org/register)

[Get a Google account](https://accounts.google.com/signup) &#x20;

[Get a Globus ID](https://globusid.org/create)
{% endhint %}

3\. Sign into the account you are linking to KBase. For ORCiD, first accept all cookies and then login with either your 16-digit ORCID iD.

<div align="left"><img src="/files/m0YJFotiOJl4X0r2e2hn" alt="" width="375"></div>

4\. The KBase account creation form will prompt you to enter some required information, which includes choosing a username and identifying your organization and department.

![](/files/9wD10q4S4RzQHHN6hy9J)

5\. You must now agree to the [KBase User Agreement](https://www.kbase.us/about/terms-and-conditions-v2/#use_agreement) and [KBase Data Policy](https://www.kbase.us/about/terms-and-conditions-v2/#data_policy).&#x20;

Read the policies and select the checkboxes that say “I have read and agree to this policy”

Once you’ve checked both boxes for User Agreement and Data Policy, you will be able to click the blue button at the bottom that says “Create KBase Account”.

![](/files/q8mlIRh7a5WKrxBwrMCT)

6\. You succesfully created a KBase account!\
\
Continue reading the rest of the documentation to learn more about the features of KBase.

If you encounter any difficulties or have any questions, please feel free to [contact us](https://www.kbase.us/support/).


# Authentication Update

KBase has simplified the sign up and sign in process for new users. You will be able to sign into KBase with a Google, Globus, or ORCiD once accounts have been linked.

{% hint style="info" %}

### New users: creating a new KBase account

You can ignore the information on this page. For details about setting up a new account, see our [step-by-step signup guide](/getting-started/sign-up/step-by-step)
{% endhint %}

### Transitioning an existing KBase account

If you already had a KBase account prior to June 9, 2017, then the first time you want to sign into KBase after the authentication change, you will need to go through a few steps to convert your existing KBase (through Globus) account to the new system. After going through the conversion, you’ll be able to sign in to KBase with your existing account or your Google account, if you choose to link your KBase account. To see what you’ll need to do, continue reading below.

When you go to [KBase](https://narrative.kbase.us/) to sign in, you will see this:

<img src="/files/-MguzQ2Y4arrlDBd4muR" alt="" width="563">

Click the Sign In button below “Use KBase” and sign-in options will appear

**Sign in using Globus**. Globus is the home of the existing KBase user accounts, and your Globus ID is your KBase ID.

When you click “[Sign in with Globus](https://www.globusid.org/login)”, a page to Log in with your Globus credentials will open:

Click Continue. You will reach another page at Globus, requesting your Globus username and password (remember, this is the same as your existing KBase account).

![](/files/-M6q79ifjOhlUJNgwxbI)

Enter your usual KBase username and password. Your username is not your email address – it should not contain an “@” sign. **If you forgot your password, use the “Forgot password?” link at the lower right.**

![](/files/-LvNoGqs4YWeqlGo-RZ0)

After entering your username and password, click the “Log In” button. You will be taken to a page where KBase requests access to your email address and Globus identities. Click the blue “Allow” button. We won’t share your email address or personal information outside of KBase.

![](/files/-LvNoX68CkOT_U8r0nlf)

You must now agree to the KBase User Agreement and KBase Data Policy. (The User Agreement and Data Policy have not changed since you last agreed to them. This is merely a reconfirmation required by the new authentication process.)

Read the [KBase User Agreement](https://www.kbase.us/about/terms-and-conditions-v2/#use_agreement) and [KBase Data Policy ](https://www.kbase.us/about/terms-and-conditions-v2/#data_policy)and check the checkboxes that say “I have read and agree to this policy”

<img src="/files/-LvNobP1Iy77bf7J6PZL" alt="" width="563">

Once you’ve checked both of the checkboxes, you will be able to click the blue button at the bottom that says “Continue to the KBase account \[username]”.

<img src="/files/-LvNov-M6WbTZagXSGVz" alt="" width="563">

Once you click that Continue button, you’re done – you should see the KBase User Interface landing page.&#x20;

Note – the system changes mean that your profile may look a bit bare-bones. Please fill in the information in your profile when you get a chance, and choose or upload a new avatar if you like.

### The next time you sign in with your Globus account

The first time you sign into KBase after the authentication change, you will need to go through all the steps described above. Subsequently, if you choose to sign-in using your Google credentials, the process will look like this:

1\. Go to [narrative.kbase.us](https://narrative.kbase.us/)

<img src="/files/-MguzQ2Y4arrlDBd4muR" alt="" width="563">

2\. Click the Sign In button below “Use KBase” and more options will appear:

3\. Click “Sign in with Globus”.&#x20;

Most of the time, you will not need to sign into Globus (as it will remember you unless you sign out of KBase), so you will see a screen like this:

![](/files/-LvNqyuWEYMdJtXlca01)

4\. Click the blue button at the bottom that says “Sign in to KBase account \[username]” and you will be at your KBase landing page!

Note that if you had signed out completely from Globus after the last time you used KBase, you might have to go through two extra steps to sign-in to Globus. Click "Continue" and then proceed to enter Globus Log in information.&#x20;

![](/files/-M6q7XoJpt5wF0aumP60)

If you encounter any difficulties or have any questions, please feel free to [contact us](https://www.kbase.us/support/).&#x20;


# Supported Browsers

The KBase website and Narrative Interface should work on most modern web browsers, but we primarily focus on supporting **Chrome**.

## Supported Browsers

The following desktop web browsers have been tested against the KBase web site and Narrative user interface. If you experience any issue using these browsers with our site please report it through the [Help Board](https://kbase-jira.atlassian.net/).&#x20;

<table data-header-hidden><thead><tr><th width="187" align="center">OS</th><th>Supported Browsers</th></tr></thead><tbody><tr><td align="center">OS</td><td>Supported Browsers</td></tr><tr><td align="center">Mac</td><td><p>Chrome</p><p>Firefox</p><p>Safari 10 (macOS Sierra) and up</p></td></tr><tr><td align="center">Windows</td><td><p>Firefox</p><p>Microsoft Edge - not recommended</p></td></tr><tr><td align="center">Linux</td><td><p>Chrome</p><p>Firefox</p></td></tr></tbody></table>

{% hint style="warning" %}
**Ad blockers may prevent KBase from working in your web browser**

Some ad-blockers and browser extensions may cause errors within the Narrative Interface. We are working to identify issues with commonly used browser extensions, but you may want to disable browser extensions if you encounter issues within KBase.<br>
{% endhint %}

## Mobile

KBase is currently not supported on mobile browsers. The KBase website and services should work with tablet browsers on iPad and Android devices in landscape orientation, but we do not currently test on these platforms so there may be compatibility issues. Watch for future updates as KBase goes responsive and mobile.

## Installation of Supported Browsers

### Chrome

Chrome is the best browser to use for KBase. The Chrome browser is produced by Google, Inc., and is freely available. Chrome operates on a rolling release, meaning that it is updated frequently and it is not easy to track specific versions. Our best advice is to make sure you are using the most recent version available. On Mac and Windows this is built into the Chrome application itself.

* [Download Chrome](http://www.google.com/chrome)

### Firefox

Firefox, produced by the non-profit Mozilla Corporation, is a freely available, open source browser for Mac, Windows, Linux, and other systems. Like Chrome, Firefox is updated frequently, so it is not feasible for us to recommend a specific (or even minimum) version. Our best advice is to make sure you are using the most recent version available. Recent versions of Firefox will automatically update themselves on Mac and Windows.

* [Download Firefox](https://www.mozilla.org/en-US/firefox/new)

### Safari

Safari is distributed with Mac OS X. (We do not support the Windows version of Safari.) Specific versions of Mac OS X support specific versions of Safari – see this [Wikipedia entry](http://en.wikipedia.org/wiki/Safari_version_history) for details.

### Microsoft Edge

[Microsoft Edge](https://www.microsoft.com/en-us/windows/microsoft-edge) is the new web browser for Windows 10; it replaces Internet Explorer. You should be able to use KBase on Edge, but this has not been extensively tested.

## Unsupported Browsers

**KBase is not supported on Internet Explorer**. It should work on Microsoft Edge, but it has not been extensively tested. For the best chance of success on Edge, be sure you are using the most recent version.

Older versions of supported browsers are also not supported. Please ensure that your browser and system are up-to-date before reporting a problem with a supported browser.

Other modern browsers, such as Opera, may work with KBase, but as we do not test with them we cannot make any assurances about their operation.


# Narrative Quick Start

A fast, one-page introduction to KBase's narrative interface

{% embed url="<https://youtu.be/WKKKXg55hs0>" %}
Narrative Quick Start Guide
{% endembed %}

## **What is a Narrative?**

In KBase, you can create shareable, reproducible workflows called **Narratives** that include data, analysis steps, results, visualizations, and commentary. This Quick Start provides an overview of how to create and use the features of Narratives.

### **Step 1. Sign up for a KBase user account**

To begin using KBase, you will first need to [Sign up for a KBase Account](/getting-started/sign-up#signing-up). You can now use your existing Google or Globus username and password to get a free KBase account.

### **Step 2. Sign in to KBase**

After you sign in, you will be taken to your Narratives (details of which can be found [here](https://kbase.us/narrative-guide/your-dashboard/)). From Narratives, you can open your existing Narratives, access others that have been shared with you, and create new Narratives. These Narratives can be accessed at [narrative.kbase.us](https://narrative.kbase.us/).

{% hint style="info" %}
If you are a new user, your 'My Narratives' and 'Shared With Me' will likely be empty, since you haven’t created any Narratives yet.
{% endhint %}

![Navigating the Navigator](/files/mP7DwL0b7IxinJV0O9Ip)

### **Step 3. Create a new Narrative**

Click the “+ New Narrative” button in Narratives to open a new untitled Narrative. A “Welcome to the Narrative Interface” box offers information and links to documentation for adding content to your Narrative. You can collapse or delete this box using the “…” dropdown menu in the top right corner.

<figure><img src="/files/xVqoIpVuQep30sXQdBNY" alt=""><figcaption><p>Welcome Cell Options</p></figcaption></figure>

**Try the Narrative tour!**

Select the “Narrative Tour” from the Help menu. This new feature walks you through the user interface, pointing out various useful aspects of it. We recommend taking the tour even if you’ve used KBase before, as it will help familiarize you with the current version.

### **Step 4. Find data to analyze**

To load data to your Narrative, click the “Add Data” button in the **Data Panel**, which can be found under the Analyze tab near the top left of the Narrative. The [Data Browser](/getting-started/narrative/explore-data) will slide out with tabs that show several data sources.

<figure><img src="/files/HYCzKOB4dsbXSgJLB1sP" alt="" width="339"><figcaption></figcaption></figure>

You can choose a dataset already in KBase. The *Example* tab, for instance, contains sample data from KBase’s reference data collection that can be used to try out the apps.

The *Import* tab allows you to upload your own data. KBase currently supports upload of a variety of data types. You can upload data up to 2GB using drag & drop or [Globus](/data/globus) for larger files. Please see the [Data Upload/Download Guide](/data/upload-download-guide) for more information.

Note: Any data you upload to KBase is private unless you choose to share it.

### **Step 5. Add data to your Narrative**

When you hover over a [data object](/getting-started/narrative/explore-data) in the Data Browser, a blue “< Add” button will appear to its left. Click the button to add the data to your Narrative.

<figure><img src="/files/weYwMOQg50XyA3J24nXP" alt=""><figcaption></figcaption></figure>

Notice that the data object now appears in the Data Panel. Once you finish adding data, exit the Data Browser by clicking either the “Close” button at the bottom left of the browser window or the arrow at the top right of the Data Panel.

<figure><img src="/files/72X1KdzDGTMgtkMENqH5" alt=""><figcaption></figcaption></figure>

Don’t forget to periodically save your Narrative by clicking the “Save” button at the top right of the interface.

<div align="center"><img src="/files/-M6BqLqwo5eWUPat8h4x" alt="" width="563"></div>

### **Step 6. Choose an App**&#x20;

Once you have added data to your Narrative, you can analyze it using one or more of the KBase apps listed in the **App Panel** below your data.

<figure><img src="/files/ZUExTAPpPHzILycmXQky" alt="" width="340"><figcaption></figcaption></figure>

There are options for filtering by category, name, input type, and more. You can designate apps you like as favorites using the star icon.

To add an app to your Narrative from the App Panel, click on its name or icon. You can get more information about an app by hovering over it until “…” appears and clicking on it.

### **Step 7. Run the App**

Fill in all the required fields in your App. Note that some app fields are “smart” and can recognize which data in your Narrative are valid input for those fields. These smart fields have a pulldown list of all the data objects from your Data Panel that you can choose from as input.

![](/files/-LvNmErEWhfTV-we03Hx)

When all required fields have been filled out, click the green “Run” button near the top left of the app cell to start the analysis.

{% hint style="danger" %}
Choosing the same output name for data objects will overwrite existing data objects of the same type with that name.&#x20;
{% endhint %}

While the analysis is running, you will see the progress in the Job Status tab of the app box. When it finishes, the results will show up in the Result tab.

Please see [the KBase Apps section of the Narrative Interface User Guide](/getting-started/narrative/analyze-data) for more details on running apps.

### **Step 8. Share your Narrative**

<img src="/files/-M6qA2Bivl59SmceNGIA" alt="" width="375">

All Narratives that you create are private by default. You can [make a Narrative public or share it ](/getting-started/narrative/share)with specific collaborators by clicking the “share” button near the top right of your Narrative.

## **Next steps and additional resources**

With these basic steps, you now can begin using the Narrative Interface to create Narratives and analyze data. For a more in-depth introduction to any of the topics mentioned here, see the [Narrative User Guide](/getting-started/narrative).


# Narrative Interface User Guide

A more leisurely introduction to KBase's Narrative interface

What is in here:&#x20;

1. [Accessing the Narrative](/getting-started/narrative/access)
2. [Tour the Narrative](/getting-started/narrative/tour)
3. [Narratives Navigator ](/getting-started/narrative/navigator)
4. [The Job Browser](/getting-started/narrative/job-browser)
5. [Create a new Narrative](/getting-started/narrative/create)
6. [Exploring data in KBase](/getting-started/narrative/explore-data)
7. [Add data to a Narrative](/getting-started/narrative/add-data)
8. [Browsing analysis tools - KBase Apps](/getting-started/narrative/add-apps)
9. [Analyze data using Apps](/getting-started/narrative/analyze-data)
10. [Revising the Narrative](/getting-started/narrative/revise)
11. [Formatting Markdown in the Narrative](/getting-started/narrative/markdown)
12. [Share a Narrative](/getting-started/narrative/share)
13. [Access and copy Narratives](/getting-started/narrative/access-and-copy)
14. [Create an Org](/getting-started/narrative/orgs)


# Access the Narrative Interface

## What is the Narrative Interface?

The Narrative Interface enables researchers to design and carry out computational experiments while creating interactive, reproducible records of the data, computational steps, and thought processes underpinning their results. These records can be shared with collaborators and published as “active papers,” called **Narratives,** which allow others to repeat the computational experiment and alter parameters or input data to produce different or improved results. The Narrative Interface is built on top of the [Jupyter Notebook](http://jupyter.org/) framework.

{% hint style="info" %}
A very quick (one-page) introduction to the Narrative Interface is available via the [Narrative Quick Start](/getting-started/quick-start).
{% endhint %}

The Narrative Interface is accessible through [narrative.kbase.us](https://narrative.kbase.us/). Once you sign in, you will see the Narratives Navigator.

There are several ways to enter the Narrative Interface from KBase:

* Use the menu at the top left of the landing page and select “Narrative Interface”.
* Open a public Narrative or one that has been shared with you
* Start a new Narrative by clicking the “+ New Narrative” button

{% hint style="warning" %}

## Make sure you have the latest version

New versions of the Narrative Interface are released periodically. In most cases, KBase will automatically update to the latest version. However, if you had an older version already running, you will see a green button near the top right of your Narrative that says “New Version Available.”\
![Screen Shot 2015-02-03 at 8.16.19 PM](/files/-MF7xjL02XcgwmgG74wY)\
Click this button to update and reload the interface. You also may need to do a hard refresh (shift-reload) in your browser window.
{% endhint %}


# Tour the Narrative Interface

When you open a new Narrative, a “Welcome” cell will appear at the top, offering quick tips for getting started and links for accessing help. You can keep the Welcome cell, collapse it by clicking the "-" button in the top right corner of the cell or delete it by selecting "Delete cell" from the "..." dropdown menu.

## The Welcome Cell

When you first create a Narrative, you may notice the Markdown cell that says "Welcome to the Narrative". This is a quick link to this documentation site and shares information on any service status for kbase.us.&#x20;

<figure><img src="/files/R77VV0GEDUQY5gRpWcld" alt=""><figcaption></figcaption></figure>

Before adding data and analysis steps to your Narrative, take some time to explore a few high-level features of the Narrative Interface.

<figure><img src="/files/v3l1TzlDaVM1v4e1DJnZ" alt=""><figcaption></figcaption></figure>

## Managing Narratives and accessing resources

### **1. Top menu**

Menu options (accessed by clicking the three dark grey small horizontal lines on the upper left of the window) while in a Narrative:

<figure><img src="/files/Z117lnjJDZOpHBWuhtwt" alt=""><figcaption></figcaption></figure>

* *Narrative Interface* lets you access other [Narratives](/getting-started/narrative/navigator) from a Narrative.
* *About the Narrative* opens a window describing Narrative Interface properties (including version number) and provides a link to notes about the latest KBase release.
* *Shutdown and Restart* closes and relaunches the Narrative Interface (which can be useful if your content fails to load or the Narrative appears frozen).

The top-level menu (accessed by clicking the three blue small horizontal lines on the upper left of the window) offers additional possibilities:

* *Narrative Interface* goes to the main Narrative Interface, which allows you to create, edit, run, and share Narratives.
* *New Narrative* opens a new Narrative in a new window. &#x20;
* *JGI Search* goes to the [Data Search](https://narrative.kbase.us/#jgi-search?q=) tab that can also be reached by the Search icon in the side bar.
* *Biochem Search* goes to the Compound and Reaction search tabs.&#x20;
* *KBase Services Status* goes to [a page that describes the current version of the Narrative Interface](https://narrative.kbase.us/#about/services).
* *About* opens a page that describes the [KBase User Interface](https://narrative.kbase.us/#about) and connects to the [KBase GitHub](https://github.com/kbase).
* *Contact KBase* lets you [contact the KBase Help Desk](https://www.kbase.us/support/) with your question, bug report, or feature suggestion.
* *Support* brings you to the [Narrative User Guide](/getting-started/narrative).

### **2. KBase logo**&#x20;

Links to the KBase home page: [kbase.us](/kbase.us).

### **3. Narrative title and creator**

&#x20;Click in this space to name (or rename) your Narrative.

### **4. Toggle view-only mode**&#x20;

In [view-only mode](/getting-started/narrative/access-and-copy), you can look at the Narrative but not run or change it.

### **5. Narrative controls**

* *Help* menu (“?”) — links to documentation as well as to the Narrative Tour.
* *Kernel* menu (intended for developers; do not use).
* *Share* button — make your Narrative public or give specific KBase users the ability to view, edit, or share it further.
* *Save* icon — Be sure to save your work often.
* *User* icon — sign in/out or access your user profile.

### **6. Narrative tabs**

* *Analyze* tab — Consists of the Data Panel (see #8) and the Apps Panel below it (see #11).&#x20;
* *Narratives* tab — Allows you to create a new Narrative, copy an existing Narrative, and see a list of other Narratives created by or shared with you.
* *Outline* tab — Shows the contents of the Narrative in a table of contents format to view Markdown cells, apps, and data objects. You can use it to navigate to different sections of the Narrative by clicking on the particular workflow cells rather than scrolling.&#x20;

## Adding data and apps

### **7.** **Data Panel** &#x20;

You can add data to your Narrative by searching KBase’s reference collection, importing datasets from external sources, or generating new data from KBase analyses. All data objects from these activities will be listed in this panel, which is empty when you create or open a new Narrative.

### **8. Data Panel controls**&#x20;

Use these icons to search, filter, sort, and refresh the data in your panel. The right arrow will open the Data Browser (see [#id-10.-add-data](#id-10.-add-data "mention")).

### **9. Data Object**&#x20;

Each data object in the panel can be expanded to see more information about the data and several options for downloading, viewing data provenance, and more. The “[Add Data to your Narrative](/getting-started/narrative/add-data)” section describes how to get more information about each data object.

### **10.** **Add Data**&#x20;

This opens the Data browser, described in detail in the “[Explore Data](/getting-started/narrative/explore-data)” and “[Add Data to your Narrative](/getting-started/narrative/add-data)” sections, provides options for working with your own data or data from KBase’s reference collection.

### **11. Apps Panel**&#x20;

This panel lists KBase analysis functions. Clicking the app name or icon beside it will add an analysis cell to your Narrative (see “[Analyze Your Data Using KBase Apps](/getting-started/narrative/analyze-data)” for details).

### **12. App Controls**&#x20;

Use these icons to search and refresh the list of apps, as well as expose apps that are available but still in beta. The right arrow will open the [App Catalog](https://narrative.kbase.us/#appcatalog), which provides advanced options for browsing apps; designating them as favorites; and filtering them based on analysis type, popularity, and more. For additional information, see [Browse KBase Analysis Tools](/getting-started/narrative/add-apps).

## Running analyses and adding commentary

### **13. Markdown Cell**

Use [Markdown language](https://blog.ghost.org/markdown/) to add rich text commentary and graphics to your Narrative.

### 14. Data Cell <a href="#id-14.-data-cell" id="id-14.-data-cell"></a>

These cells share details and describe individual data objects present within the Narrative. There are cell controls at the top right, similar to app and markdown cells.&#x20;

### **15.** **App Cell**&#x20;

There are several types of cells, including those displaying app inputs and outputs, text and commentary, and code cells for writing custom scripts. The cell controls are at the top right of each cell.

### **16.** **Markdown and Code Cell buttons**&#x20;

Click to add these types of cells to your Narrative. Stay tuned for more documentation on how to use code cells in the future!


# Narrative Navigator

Welcome to Narratives! You may notice there are a few changes when you sign in to narrative.kbase.us. Let's get you up to speed with a tour.&#x20;

{% embed url="<https://youtu.be/Nb5lTlB0NBo>" %}

<figure><img src="/files/7s6nkOB2njRWAKH669tX" alt=""><figcaption><p>Overview of Narrative Navigator Tabs</p></figcaption></figure>

* There are five tabs for exploring and locating existing Narratives: **My Narratives, Shared with Me, Tutorials,** and **Public.**
* The most recently updated Narratives appear first, but can be searched or sorted using the dropdown menu - *Recently updated, Least recently updated, Recently created, Oldest, Lexographic (A-Z), Reverse Lexicographic (Z-A)*.
* A list of available Narratives will be on the left panel and information on the selected Narrative on the right panel. Here you can see an overview of the Data within the Narrative and a Preview. &#x20;
  * Icons under a Narrative name indicate which apps the Narrative includes, how many markdown and code cells it contains, the number of times the Narrative has been shared, whether there are any running jobs, total running time for all completed jobs, and more.
* Click on the selected Narrative name to open it.
* Click on the *Gear* icon next to the Narrative title for management options to:&#x20;
  * Manage Sharing
  * Copy this Narrative
  * Rename
  * Link to Organization
  * Delete

<figure><img src="/files/xM3X1ASORNCeDr8lsWuh" alt="" width="208"><figcaption></figcaption></figure>

* You can use the *Search Box* at the top of each panel to find a particular Narrative within that category.
* Click the *+ New Narrative* button at the top to create a new Narrative.

<figure><img src="/files/S4qKQsuWTYzMK4PHAnaG" alt=""><figcaption><p>Click on the + New Narrative button on the right side of the top bar to create a new Narrative</p></figcaption></figure>

* Note that **Public Narratives** have not necessarily been reviewed or tested by the KBase team; they are simply Narratives that have been made public by their owners.

## Where can I go from here?

The Narrative Navigator offers many launch points. To the left of the main panel you will see seven icons allowing you to toggle between **Navigator**, **Orgs**, **Catalog**, **Search**, **Jobs**, **Account,** and **Feed** tabs.

<figure><img src="/files/uOSStKfxq7bguYINvDCD" alt="" width="72"><figcaption></figcaption></figure>

### **Navigator**

* Takes you to your accessible [**Narratives**](/getting-started/narrative/create) through the **Narrative Navigator** and features described above.

### **Orgs**

* Takes you to a list of the [**Organizations**](/getting-started/narrative/orgs) within KBase, along with descriptions and statistics for the Organizations, including Narratives and Apps.&#x20;
* The default shows "My Orgs," the Organizations you belong to as a member.&#x20;
* You can view and search all Organizations by selecting the toggle button on the right hand side of the page for "All Orgs."&#x20;
* Under "All Orgs" you can search and select different Organizations and request to join by clicking on the "Join this Organization" button.&#x20;

### **Catalog**

* You can use the [**App Catalog**](/apps/analysis) to browse, search for, and see the specifications of each app KBase has to offer.

### **Search**

* The **Data Search** tab gives you the ability to search JGI data (*Beta*) as well as KBase user and reference data.
* It allows you to search for data present in your own Narratives, those shared with you, and those made public by their owners.
* It also enables you to search for reads and assemblies contained in the Joint Genome Institute (JGI) Genomes Online Database (GOLD) and add them directly to your Narratives.

### **Jobs**

* The [**Job Browser**](/getting-started/narrative/job-browser) allows you to view apps that you have run recently, see their progress, search for specific runs, and more.
* You can filter the runs by *Finished*, *Queued*, *Running*, *Success*, and *Error*
* Query results also display basic information surround the runs such as the Narrative they can be found in, App ID, Submission Time, Queue Time, Run Time, and Status

### **Account**

* In the [**Account Manager**](/manage-account) you have access to edit your profile information.&#x20;
* [Link additional accounts](/manage-account/link-accounts), including Google, Globus, and ORCiD [<img src="/files/-M9K9Aepzk1FNZ3jYzOv" alt="" data-size="line">](https://orcid.org/) .&#x20;
* From this tab you can also get to your **Profile Page** which gives an overview of your Narratives and your Collaborators.
* You can access the user profile of any of your Collaborators by simply clicking on the name under the **Your Collaborators** panel while in your **Profile Page.**

### **Feeds**

* Takes you to the **Notification Feeds** page to view notifications from KBase.&#x20;


# Create a Narrative

## Creating a new Narrative

You can create a new Narrative from your [Narratives](/getting-started/narrative/navigator) by clicking the “+ New Narrative” button, which will take you to a new Narrative Interface.&#x20;

<figure><img src="/files/oFTrIMrSKe0SuwZzl9cO" alt=""><figcaption></figcaption></figure>

If you are already in the Narrative Interface, you can start a new Narrative by selecting the *Narratives* tab (between *Analyze* and *Outline*) and clicking the button labeled “+ New Narrative.”

<figure><img src="/files/dXTfKWvJM6Xlw4yj9HKC" alt=""><figcaption></figcaption></figure>

The new Narrative will open in a new browser tab and contain a default “Welcome to the Narrative” cell that provides brief instructions for using the Narrative Interface. You can retain this cell or delete it with the “Delete cell” option in the “…” dropdown menu in the upper right corner of the cell.

You can name the Narrative now by clicking on “Untitled” and entering a new name.

Note: The new Narrative will not appear in your Narratives list until it is named something other than ‘Untitled’ and you have clicked the Save button to save it.

We hope that after working on your Narrative, you will want to [share it](/getting-started/narrative/share) with other users. **Any Narrative you create is private until you choose to share it**.

In the sections that follow, we will explain how to [explore data](/getting-started/narrative/explore-data), [add data to your Narrative](/getting-started/narrative/add-data), and use [apps](/getting-started/narrative/add-apps) to analyze it.

## Copy and run Narratives created by others

If you don’t feel quite ready to create your own Narrative, you can copy and run a Narrative that someone else has shared with you. Please see the [Access and Copy Narratives](/getting-started/narrative/access-and-copy) section for more information.


# Explore Data

There are several ways to find data and add it to your Narrative. You can select data available within KBase, [upload your own files, or import datasets from external resources](/getting-started/narrative/add-data) such as the DOE Joint Genome Institute (JGI) or NCBI. This section will describe how to explore data already in KBase. Instructions for adding data to your Narrative and importing external files are provided in the next section “[Add Data to Your Narrative](/getting-started/narrative/add-data).”

## The Data Browser

There are several ways to explore the [wide range of data available in KBase](/data).

The Data Browser slide-out is accessed by clicking the right-facing arrow in the menu or the red rectangular "Add Data" button in the Data Pane.

<figure><img src="/files/4svRH8UgofoElhlp0HRk" alt="" width="339"><figcaption><p>Emply Data Pane</p></figcaption></figure>

* **My Data** — data you have already loaded or analyzed in another Narrative
* **Shared With Me** — data included in Narratives that have been shared with you
* **Public** — data in KBase that is accessible to everyone
* **Example** — example datasets that can be used as inputs to apps and methods
* **Import** — mechanism allowing you to import your own data (of supported types) to your Narrative.

<figure><img src="/files/yFaVfv5F8u4FM18dvuKL" alt=""><figcaption><p>Accessing Data through the Data pane</p></figcaption></figure>

Find a public reference genome available under the *Public* tab to examine information and metadata about that genome.

Clicking the *Public* tab displays a list of data objects available in the KBase reference collection.

<figure><img src="/files/22Au260e1vdgg1alzsK9" alt="" width="563"><figcaption></figcaption></figure>

Use the Dropdown tab to select the Public Database to filter the displayed data objects.&#x20;

<figure><img src="/files/PjsIzU5MTKs5Pm4pcwh7" alt="" width="201"><figcaption></figcaption></figure>

You can also use the “*Filter data..."* field in the Data Browser to find data objects whose names include the text you’ve typed in the search box. (Note — searches are not case-sensitive).&#x20;

If you hover over a data object in the Data Browser panel, three small images appear: an *"< Add " button* in a blue rectangle to the left of the object name and *binoculars* icon and a *graph-like* icon to the right. These latter two icons open a **Data Landing** page and a **Provenance** page, respectively.

<figure><img src="/files/jIfejFtw9ZVfVCEOLnMu" alt="" width="563"><figcaption></figcaption></figure>

Go ahead and click the "< Add" button to add the data object to your Narrative. Notice that the data now appears in the Narrative Data Pane.&#x20;

<figure><img src="/files/L0AHb5UIjZP1X2DcrLvd" alt="" width="563"><figcaption></figcaption></figure>

## Information in the Data Pane

The Data Pane shows all the data that you’ve added to your Narrative. (The [next ](/getting-started/narrative/add-data)[section](/getting-started/narrative/add-data) of this guide discusses in more detail how to add data to your Narrative.)

You should have at least one object in your Data Pane. As you add or generate more data during the course of your analyses, the number of objects in this panel will increase. You can search, sort, or filter the list using the icons in the Data Panel header.

You can access more details about a particular data object by hovering over the object and clicking the ”...” that appears in the right-hand side of the data object.

<img src="/files/-M6Vp7I84gZxOPmBySAb" alt="" width="563">

The expanded view of the data object reveals icons that let you examine or manage the data.

* The **arrow in** icon for **"Show Apps with this as input"** shows the apps (in the App Panel) that can use this type of data for input.
* The **arrow out** icon for the **"Show Apps with this as output"** button that shows the apps (in the App Panel) that will generate this type of data as a data object.&#x20;
* The **binoculars** icon is the **"Explore data"** button that open a Data Landing page for the data object.
* The **page with three lines and folded upper right corner** icon is the **"View associated report"** button and will open the Data View report page for the data object.&#x20;
* The **counter clockwise circle with clock hands** is the **"View history to revert changes"** icon allows you to see and revert to previous versions of the data.
* The **graph-like** icon is the **"View data provenance and relationships"** button that opens the Provenance page (described below).
* The **download** icon is the **"Export/Download data"** button that lets you download a data object to your local computer. (Note — This capability is still in development; most data objects currently can be downloaded only in the JSON format. In the future, you will be able to choose from a variety of common formats such as GenBank.)
* The **“A”** icon is the **"Rename data"** button that lets you rename your data object. Use with caution. If a Narrative was already using the old object name in analyses, they might stop working. Object names can contain only letters, digits, and underscores; no spaces or other special characters are allowed.
* The **trash can** icon is the **"Delete data"** button that lets you delete a data object from your Narrative. (If the data came from a public or shared data source, only the copy within your Narrative will be deleted.)

## Data viewers

Many types of data in KBase have viewers that allow you to learn more about the data. These viewers can be accessed two ways from the Data Panel:

1. Click the name of a data object, and its viewer will be added to your Narrative.
2. Drag the object from the Data Panel and drop it into the main part of your Narrative.

Below is a data cell for the a genome that is within the Data Pane.

<figure><img src="/files/GRK5DzegCPKxZ4bbSxxn" alt="" width="563"><figcaption></figcaption></figure>

Notice that this genome viewer has tabs for an overview (including GC content, taxonomy information, size, and more) and a list of contigs and genes. Each contig and gene entry in these lists is clickable, opening either a contig browser or a tab with expanded information about the gene.

<figure><img src="/files/bNc32YNPy4h4dudoJOI6" alt="" width="563"><figcaption></figcaption></figure>

**Sorting table entries in the viewer**

You can sort the table entries by clicking on a column header to sort by that field (e.g., Length). Clicking the same column header again will reverse the sort order. For example, the screenshots below show the table sorted in descending order by *Contig name* and then in descending order by *Length.*

You can also sort by more than one column at a time by clicking one column header and then Shift-clicking other column headers. For example, here we have sorted in ascending order by contig length, and then (by shift-clicking the Genes column header) in ascending or descending order by number of genes.

![](/files/-LvNyZFgfxvxVuUxT5NY)

There are many types of viewers in KBase in addition to the Genome viewer discussed here. Different viewers may have different options to explore.

To remove a viewer from your Narrative, click the trashcan icon in the top right of the viewer cell. You can always re-add the viewer; removing it from your Narrative doesn’t delete the data object itself.

## Data Landing and Provenance pages

Data Landing (also known as *Data Summary*) and Provenance pages are ways to find out even more information about a data object. You can access these pages using the icons that appear when you hover over a data object in the Data Browser or click on one of the objects in your Data Panel. The binoculars icon opens a new Data Landing page about that particular data object, while the graph-like icon opens a Provenance page.

**Data Landing pages** (which are still in development) provide both known and contextual information about a data object, allowing users to examine various particulars about the data and, eventually, compare it to other data objects. Depending on the type of data and where it came from, different sorts of information might be presented. For example, for a genome already in KBase (such as the one we loaded earlier, [Vibrio brasiliensis LMG 20546](https://narrative.kbase.us#dataview/KBasePublicGenomesV4/kb%7Cg.3791)), you will see:

* A data object summary
* A data provenance and reference network (collapsed by default; open it by clicking the > on the left)
* A genome overview panel that lists information about the genome, including its biological domain, DNA length, etc.
* A description of the species (including photo, if available)
* Publications that mention this genome
* The taxonomic lineage
* A species tree (with the option to create a tree if one doesn’t yet exist for this species)
* A contig browser
* Functional categories
* and more

**Provenance pages** are one way to facilitate reproducibility and transparency of scientific results, two key principles of KBase design. All data in KBase is versioned, and older versions of a data object can be accessed from the Provenance page. This page records and illustrates how data is derived and modified in KBase, including how it entered the system, whether it was produced through analysis of other objects, who generated the data, and when. You also can identify the original “owners” of the data, allowing you to contact them or see their shared analyses.

The image below shows the Provenance page for a newly imported genome.

<figure><img src="/files/dnLJDBMlPY5iSdPgwnJG" alt=""><figcaption></figcaption></figure>

In the [next section](/getting-started/narrative/add-data) of this guide, we will discuss in more detail how to add data to your Narrative.


# Add Data to Your Narrative

Now that you are familiar with ways to find and explore data in KBase, you can select or upload data to analyze. The **Data Panel** in a Narrative shows the data objects that are currently available in that particular Narrative.&#x20;

<figure><img src="/files/Zi95U5zhKLDa2ycF2GcF" alt=""><figcaption></figcaption></figure>

From the Data Panel, you can access the data slide-out, which allows you to search for data of interest and add it to your Narrative. In the Data Panel, click the "Add Data" button, the “+” button, or the arrow at the upper right of the panel to access the **Data Browser** slide-out.

<figure><img src="/files/yFaVfv5F8u4FM18dvuKL" alt=""><figcaption><p>Add and explore data to your Narrative. </p></figcaption></figure>

{% hint style="info" %}
**Data Privacy**\
Any data that you upload to KBase is kept private unless you explicitly choose to share it. You can share any of your Narratives (including their associated data) with one or more specific users, or make it publicly available to all KBase users. Please see the [Sharing](/getting-started/narrative/share) page for more information about how to do that. The [Terms and Conditions page](https://www.kbase.us/about/terms-and-conditions-v2/#data_policy) describes the KBase data policy.
{% endhint %}

The first four tabs of the **Data Browser** (*My Data*, *Shared With Me*, *Public*, and *Example*) let you search data that is already in KBase. The *Import* tab lets you import data from your computer to your Narrative so that you can analyze it in KBase.

* The *My Data* tab shows data objects that you have added to your project. You may need to refresh this tab to see your most recently added data.
* The *Shared With Me* and *Public* tabs display datasets that others have loaded and made accessible to you (or to everyone). Data within each group is searchable and can be filtered. Since there are a large number of public datasets, you may wish to filter them by data type (using the pulldown selector on the left) or narrow the list by searching (in the “Search data” text box) for specific text in the data objects’ names.
* The *Example* tab shows datasets that have been pre-loaded for use with particular apps. These can be handy for trying out the Narrative Interface.
* The *Import* tab allows you to upload your own datasets for analysis.

## Data available in KBase

If you hover your cursor over any data object under the first four tabs, options will appear allowing you to add that object to the Narrative or find out more about it.

<figure><img src="/files/ZAFLfy2QRFebTTodYTaN" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/AI5XeeXcviMvYHSv7MZa" alt=""><figcaption></figcaption></figure>

In the [previous section](/getting-started/narrative/explore-data), we described the process of adding a genome to your Narrative from the public data in KBase. Now let’s check out the different data types available under the *Example* tab. The icon to the left of each data object represents its data type.

<figure><img src="/files/6mspWcBJex6zJz2UQTcU" alt="" width="494"><figcaption></figcaption></figure>

As described in the previous section, the blue "< Add" button next to these icons lets you add the data object to your Narrative. You can add more more data to your Narrative from the *My Data*, *Shared with Me*, *Public*, and *Example* tabs to try out with various KBase Apps.

![](/files/GQzlHw0Q9Hk21vWHnMGn)

## Uploading data from external sources

The *Import* tab lets you drag & drop data from your computer into your Staging Area to import into your Narrative, where you can then analyze it using KBase’s analysis apps.

To upload data from your computer (or a [Globus endpoint](/data/globus) or URL), choose the rightmost tab of the Data Browser to open the *Import* tab.

<figure><img src="/files/NYoEv52okOSubdTSxyns" alt=""><figcaption></figcaption></figure>

You can then click the “?” icon just below the drop zone to the right to launch a short interactive tour that shows the different parts of this user interface.

Getting data from your computer to your KBase Narrative is a three-step process:

1. Drag and drop the data file(s) from your computer to the new Import tab to upload them to your Staging Area.
2. Choose a format for importing data from your Staging Area into your Narrative.
3. Run the Import app that is created.

{% hint style="info" %}
**What is a Staging Area?**\
Your Staging Area is a temporary holding place for your uploaded data files. It is private to you–no one else can see the data in your Staging Area. The files in your Staging Area stay there and can be seen from any of your Narratives. When you’re ready to add data to a Narrative, you can choose a data type from the pulldown menu next to a data file in your staging area in order to add it to your Narrative as an object of that type–see below for instructions.
{% endhint %}

{% hint style="info" %}
**Why have a Staging Area?**\
It’s more robust and extensible this way. Unlike our previous importer, the staging upload can handle large data files without timing out. The user interface is more intuitive: you can drag and drop one or more data files into your Staging Area, or even whole directories or zip files. Finally, the new importer is easier for our developers to extend to new data types.
{% endhint %}

{% hint style="info" %}
**What's the difference between 'Upload' and 'Import'?**\
In this documentation, we will use “Upload” to refer to getting data from your computer into your Staging Area and “Import” to refer to the process of converting a data file from your Staging Area into a data object in your Narrative that you can analyze with any of our analysis apps. Sometimes we use the term “Import” to cover the Upload+Import process.
{% endhint %}

## Drag & Drop data files into your Staging Area

{% hint style="warning" %}
**Drag & Drop Limitations**\
The drag & drop from your local computer works for many files, but there is a size limit that depends on your computer and browser. We recommend using [Globus Online transfer](/data/globus) for files over 2GB.
{% endhint %}

Find the file(s) you want to import into your KBase account, and drag them into the drop zone (the rectangular area surrounded by a dashed line). You can select multiple files from your computer and drag them all at once. (In the example below, the user is dragging two files into the dashed area.) You can also select a folder of data files and drag the folder into the Staging Area drop zone.

![](/files/HY3XAFNuguG6aflkVzdX)

If you don’t like using drag & drop, you can instead click in the upload area to open a file chooser and select a file from your computer to upload.

While files are being uploaded to your Staging Area, you’ll see a green progress bar.

![](/files/-LvRXAR-QaD5McNW_CVD)

When the file is done uploading, you will see it appear in the list of files in your Staging Area. If you don’t see your file, try clicking the refresh button.

![Staging Area](/files/fygS1pYVo223C9qaXAKi)

By default, these are sorted by age, with the most recently uploaded file at the top. To sort the list by other fields, such as name or size, click a column header.

{% hint style="warning" %}
**90-day lifetime for files in your Staging Area**\
Your Staging Area is meant to be a temporary holding area for data you want to import into your KBase account. After adding files to your Staging Area, be sure to import them into your Narrative soon, as files in the Staging Area are automatically removed after 90 days. Data objects imported into a Narrative will exist indefinitely.
{% endhint %}

## Transferring data from a Globus endpoint

Globus is a data management and file transfer system that can facilitate bulk transfer of data (either large data files or a large number of files) between two endpoints. The endpoints that apply here are KBase, JGI, and your local computer. The KBase endpoint is called “KBase Bulk Share,” and JGI has their own way to link to Globus. To do any transfer using Globus, you will [need a Globus account](https://www.globusid.org/create).

<figure><img src="/files/ycgwBuO4Iw9gWMEs2U4D" alt=""><figcaption></figcaption></figure>

See [Transferring Data with Globus](/data/globus) for more documentation on using Globus.

{% hint style="info" %}
**Uploading data from JGI**\
If you are a JGI user, you can transfer public genome reads and assemblies (as well as your private data and annotated genomes) from JGI to your KBase account—see [the JGI data transfer page](/data/jgi-transfer) for instructions.
{% endhint %}

## Uploading data from a URL

Below the link to Globus, another link says “Click here to use an App to upload from a public URL” (for example, a GenBank ftp URL, or a Dropbox or Google Drive URL that is publicly accessible).

<figure><img src="/files/pPZtcrqLLzUrqcnLEi6k" alt=""><figcaption></figcaption></figure>

Clicking this link adds the “Upload File to Staging from Web” App to your Narrative:

![](/files/-LvRXaE-kAgy0ldEfi8G)

There are also several apps that import specific file types (single- or paired-end reads or SRA files) from a URL directly to your Narrative, bypassing your Staging Area. These are available from the Apps Panel and the [App Catalog](https://narrative.kbase.us/#appcatalog).

## Importing files from the Staging Area to the KBase Narrative

### Import Basics

The files in your Staging Area are ready to import into your Narrative as KBase data objects that can be used in your analyses. To import a file from your Staging Area, choose a format (data type) from the pulldown menu to the right of the file’s age. (You can find out more about KBase data types and accepted formats in the [Upload/Download Guide](/data/upload-download-guide).) Then click the import icon to the right of the format menu.

![](/files/-LvRXgO337L0Vr6_j_Rv)

When you click the import icon, the Data Browser slides shut and an Import App cell for each data type or a Bulk Import App cell will be created in your Narrative, with the appropriate parameters filled in. For example, here’s an import app created by choosing “GenBank” as the import format:

![](/files/-LvRXl3Fk8ng6DcckkJD)

{% hint style="info" %}
**Importing different types of data**\
This example shows how to import GenBank data. Please see the [Upload/Download Guide](/data/upload-download-guide) for detailed instructions for other supported data types and importing multiple files.
{% endhint %}

If the GenBank file came from a different source, use the pulldown menu to select it. You can change the output object name, if desired, and then click the "Run" button to start the import. When the import is done, you should see the message “Finished with success” near the top of the app cell, and some information about the app run.

![](/files/-LvRXpZfFjcglRhw3RnG)

If you look at your Data Panel, you should see the new data object created by the import.<br>

![](/files/-LvRYAdThqsZrwspu1St)

You can now use this data object as input into the relevant KBase apps. If you want to see which apps accept a particular data type as input, you can click the “…” menu in the data object cell that appears when you hover over it, and then use the “Show Apps with this as input” icon to filter the apps in the Apps Panel.

![](/files/-LvRYTCeEGpm7yosKQNC)

{% hint style="warning" %}
**What if my import fails?**\
Sometimes an import doesn’t work. One of the most common causes of failure is attempting to import a file that’s the wrong data type, or not the expected format for that data type. For example, the screenshot below shows what a user got when they tried to import a GenBank file as Media.

You can also look at [common import errors and their meanings.](/troubleshooting/job-errors/import)
{% endhint %}

![](/files/-LvRYXYCggAAQQd9mCg8)

If the importer objected to something in your file, check the [Data Upload/Download guide](/data/upload-download-guide) for details about the relevant format.

\
In some cases, the cause of an import error will not be obvious. If you can’t figure out why your import isn’t working, please contact us (via the [Help Board](/troubleshooting/support#contact-us)) for help. Note, however, that no one besides you has access to your Staging Area, so we will not be able to see the files you uploaded to your Staging Area. You may need to attach your input file to your Help Board ticket in order for us to diagnose the problem.

### Bulk Import for Multiple files

Multiple files can be imported together in a single Import App as part of [the bulk import feature](/data/upload-download-guide/uploads). When you select the file type from "Import As...," you will be able to check multiple files in the staging area and import them all by clicking "Import Selected."&#x20;

![Bulk Import Staging Area. You can select the appropriate type for multiple files, then "Import Selected" to import them simultaneously. ](/files/-MgGo7TNCJiyX4IKFjv_)

When you import selected, there will be a new cell type which contains a tab for each different file type you selected from the staging area. You will fill out the same parameters for each file type by clicking the file type in the "Data Types" column on the left; these are the same parameters as single imports.

![Bulk Import Cell. File types can be selected under the data type column to set the parameters for each type.](/files/-MgGs9JWLD-T81ihCKJ5)

Supported file types include:

* Assembly - FASTA
* SRA Reads&#x20;
* FASTQ Reads Interleaved - FASTQ Interleaved reads
* FASTQ Reads Noninterleaved - FASTQ paired-end reads or FASTQ single-end reads library
* GenBank Genome - GenBank
* GFF Metagenome - FASTA and GFF3

## Getting more information about files in your Staging Area

The list of files in your Staging Area includes their name, size, and age (from when they were uploaded). If you have many files in your Staging Area, use the Search box to locate specific files.

![](/files/ANwcHIjRVUexswga5Jp7)

Compressed or zipped files have a little double-arrow icon next to the filename. You can click this icon to unpack these files. For how to upload and import file types, follow [this guide](/data/upload-download-guide/compressed-files).&#x20;

For more information about a file in your staging area, click the arrow to the left of the filename to open a tab like this:

![](/files/-LvRYrMp-pL3h2HT5Z-P)

You can click “First 10 lines” or “Last 10 lines” to see that portion of the file:

![](/files/-LvRZ4F4g1mFiv9P_Soz)

Opening the information about a file in your Staging Area also reveals a trash can icon that allows you to remove the file from your staging area.\
![](/files/-LvRZDHADa4aUC3CmxA6)\
You will be asked to confirm that you want to delete the file. This action is not reversible.

\
Note that if you had already imported the file to your Narrative as a data object, that object won’t go away when you delete the file in your Staging Area. If you want to delete a data object, you can do that in your Data Panel.

{% hint style="info" %}
**In progress**\
KBase’s import functionality and user interface are under active development. We [welcome your bug reports and suggestions](/troubleshooting/report) for future upload functionality.
{% endhint %}

## Now what?

Once you have added data–your own data or reference data that is already in KBase–to your Narrative, you will be ready for the exciting part: analyzing it! The next two sections describe how to [choose](/getting-started/narrative/add-apps) and [run an app](/getting-started/narrative/analyze-data) to analyze your data.


# Browse KBase Analysis Tools

A Narrative can include any number of analysis steps that are added by selecting one or more apps from the Apps Panel directly below your data. Apps are grouped by category. Click the arrow to the left of a category name to see the individual apps in that category.

![](/files/-Lv6zVODH3GeFQlGG0jK)

You can browse KBase Apps by (1) scrolling through the list in the Apps Panel; (2) searching for an app by name; or (3) using the App Catalog (click the right arrow in the top right corner of the Apps Panel to open it) to further explore, filter, and designate apps as favorites.

## **Search for Apps in the Apps Panel**

![](/files/-LvRZvdqfTXfkYSACMtm)

Click the search icon (magnifying glass) at the top of the panel to begin a text search for apps. Once you enter a term, the list will show only the categories that have apps whose names include the text you typed. As before, click the right-arrow to the left of a category name to see matching apps in that category. (To see the complete list again, click the "x" to the right of the search box.)

![](/files/-LvR_5g-5gkfgqZQ7XXI)

Hovering over an app in the list will show the more “…” icon. Clicking the "..." icon will open an app details page featuring an expanded description of the app, along with technical information about its inputs, outputs, parameters, and more. Links to additional documentation such as tutorials or FAQs also may be provided.

<img src="/files/-MFC5ZpqPtf66LRrLMYK" alt="" width="563">

## **Access the App Catalog**

All available apps can be found in the App Catalog, which is accessible three ways:

* External [App Catalog](https://kbase.us/applist/) (no user account or sign-in required for browsing).
* From KBase by clicking on the Catalog icon in the Menu <img src="/files/rAVXgXN0K6bksrdHUj8j" alt="" data-size="line">&#x20;
* From the Narrative Interface, by clicking the right arrow at the top of the Apps Panel.

<figure><img src="/files/p1fYexd0KWMCYrf0Kizz" alt="" width="375"><figcaption></figcaption></figure>

The App Catalog panel will slide out to show available apps. (The “View App Catalog” link at the top opens the App Catalog in a separate web browser window.)

<figure><img src="/files/zMd20l0grw4P8aOo11Vz" alt=""><figcaption></figcaption></figure>

**Add your own app!**

The App Catalog is growing to contain not only the apps generated by KBase staff, but also those created and contributed by outside developers. KBase’s [Software Development Toolkit (SDK)](https://kbase.github.io/kb_sdk_docs/), provides a mechanism for users to add their own open-source, open-license tools to the system as new KBase apps.

## **Sort and Filter Apps**

[![](/files/-LvR_KMJZERZfFkBsouY)](https://kbase.us/wp-content/uploads/2014/12/Screen-Shot-2017-01-31-at-10.11.24-PM.png)

From the App Catalog, you can search for apps by name and use several options to organize and filter how they are displayed (the default sort is by category). Click the “Organize by” drop-down menu to see these options.

* **My Favorites** — Lists your favorited apps (those you’ve clicked the star icon to “favorite”) at the top of the catalog.
* **Favorites Count and Run Count** — Suggests some of the most popular and frequently used apps in KBase.
* **Name** — Filters apps by name (alphabetically or reverse alphabetically).
* **Category** (default sort) — Displays apps in groups by analysis type including annotation, assembly, communities, comparative genomics, expression, metabolic modeling, and more.
* **Module** — Sorts apps based on which SDK module they belong to.
* **Developer** — Displays apps alphabetically by developers’ KBase usernames, which are linked to their profile pages.
* **Input Types and Output Types** —Groups apps based on the type of data that they take as input or generate as output. Each data type is linked to a specification page that provides some technical details, including version, description, structure, and more. This sort can be useful for determining the data you must add to your Narrative before using a specific app. The *Example* tab of the [Data Browser](/getting-started/narrative/add-data) provides access to various data types to help you get started.

## **Explore Individual Apps**

![](/files/-LvR_gXa6N3nu_Htbqtb)![Screen Shot 2017-01-31 at 10.20.11 PM](/files/-LvR_iqFKV2xjJp-Yngk)

Similar to the App Panel, the App Catalog lets you (1) see a short description of each app (by hovering over the “i”–see second screenshot at right), (2) access the app detail page using the “more” link, and (3) add an app to your favorites by clicking the star.

The number beside the star icon indicates how many other users have selected the app as a favorite, while the next number shows how many times the app has been run.

You can access the app module (a group of related apps) by clicking its hyperlink (in this example, “fba\_tools”). Also, some apps will link to the KBase profile pages of their developers (in this case, chenry).

**Feeling adventurous?&#x20;*****Go beta!***

To try out apps still in development, click the little “R” in the Apps Panel, changing it to a “B”. This will show apps that are in Beta (i.e., still being tested and debugged). Click the “B” again to go back to “R” (show only released apps).

![](/files/-M8MCxg8OvWslFnVxHai)

Now that you know how to find an app of interest, the [next section](/getting-started/narrative/analyze-data) will discuss how to add it to your Narrative so you can begin analyzing data.


# Analyze Data Using KBase Apps

Once exploring available [KBase Apps ](https://kbase.us/applist/)and determining which ones work with the data objects in your Narrative, you can start analyzing your data.

## Add an App to your Narrative

To add a KBase analysis app, find the app of interest and click its name or the icon to the left of the app name. A box (called a *cell*) containing the chosen app will appear in the main Narrative panel.

<div align="center"><img src="/files/-MgvqNJhJ_SUmwDfLjIy" alt=""></div>

A few things to notice about the app cell:

#### **App Cells** have three tabs&#x20;

1. Configure (where you set the parameters; it’s what you see when you add a new app cell)
2. Info
3. Job Status&#x20;
4. Result

The last two tabs are filled in once you run the app; these are discussed later.

![](/files/-MFDAQBFFP-EnseqGF_x)

* The up and down arrows let you move the cell up or down in your Narrative.

![](/files/-M6BKbH09aqeEZEDKtc8)

* The App “…” dropdown menu offers several options for showing the code that will execute the app; getting more info (in a separate browser window) about the app; and deleting the cell. These are discussed more in the section called “[Revise Your Narrative](/getting-started/narrative/revise)“.

<img src="/files/-M6qBQfUT8H2KnY7TsB9" alt="" width="375">

* The "-/+" in the App Cell menu is the collapse/expand cell option. When the App Cell is expanded, the farthest right option will be a "-" to collapse the cell. When the App Cell is collapsed, the farthest right option will be a "+" to expand the cell.  You can also double click on an App Cell to expand/collapse the app.&#x20;

<img src="/files/-M6BNayU2rrcxCN03Xrd" alt="" width="375">

<img src="/files/-M6BNzi09xuLOMJ5rSot" alt="" width="375">

* Every app has some parameters (fields) that must be filled in before you can run the app.

## Fill in parameters

After you add an app to your Narrative, the required parameters must be filled in before you can run it. For most apps, you will need to select the input data object(s). Other parameters may also need to be set. A red bar to the right of a field indicates that it is a required field and you have not yet entered a valid value in it. Other indicators, such as banners for errors and warnings, may appear. Read the message and hover over any icons to reveal hints. Once you have filled or corrected the field, the indicator should disappear.

Some app fields are “smart” and know which data in your Narrative is valid for that field. These “smart” fields have a pulldown list of data objects that you can choose from. (Remember, only data (of the appropriate type) that you have already added to this particular Narrative will be shown in that list. You can access your data from other Narratives via the My Data and Shared with me tabs in the [Data Panel](/getting-started/narrative/add-data).)

Some fields are also required, but they will be pre-filled with default values. You can change their values if you chose, or leave them. Additionally, some apps have optional “advanced options” that you can reveal by clicking on the “advanced options” link at the bottom of the cell.

The green "Run" button is enabled when all required fields are filled in.

**Save Your Work**

<img src="/files/-M6qBbym18B0GyDNzgMO" alt="" width="375">

As you add to your Narrative, be sure to save your work frequently, using the Save button at the top right of the screen.

## Run the App

When you have filled in all the required fields, the *Run* button on the left side of the app will be enabled; click it to start running the app. Once clicked, the green *Run* button will turn to a red *Cancel* button. When the analysis starts running, you will see information in the *Job Status* ta&#x62;***.***

![](/files/-LvReqORphQuJX5bxoan)

Depending on the analysis, the analysis might finish within a few seconds or could take hours to run. You can save your Narrative and come back to it later–the analysis will continue running even if you don’t have the Narrative open.

**Examine and download results**

When an app finishes running, you will see a summary in the *Results* tab, as well as an output cell below your app cell. Also, if the analysis generates a new data object, that object will be added to your Data Panel.

![](/files/-LvRexfd6JwbhoYl35bB)

In the example, the “Annotate Microbial Contigs” app, once run, produced an output cell that provides information about the newly annotated *Shewanella* genome object.

This output cell has three tabs: Overview (displayed above), *Browse* Contigs, which lists the large assembled pieces of a genome, and *Browse Features*, which allows browsing of all the annotated genes.

Different apps create different types of output cells when they run, depending on which type of data object is output by the app. The genome output cell shown above is an example of a *Data Viewer.* Data Viewers are described in the [Explore Data](/getting-started/narrative/explore-data) section of this guide.

![](/files/-LvRf6bIabSz2Z4k6Dc9)

In addition to creating a new output cell, the app we ran created a new data object and added it to our Narrative. The newly annotated genome object now appears at the top of the Data Panel.

![](/files/-LvRfD-oJ74nLuCz0MFt)

Remember that you can hover over an object in the Data Panel and click the “..." that will appear to see more information and options, including an option (see image) for downloading the object in GenBank or JSON format (see the [Download Guide](/data/upload-download-guide) for more information). In the future, you will be able to choose from a wider variety of common formats, allowing you to download results to use with other tools and send to colleagues.

## **Conduct further analyses**

Once you have reviewed your results, you can use your newly generated data in additional analysis steps. Remember, you can click on the data object in the Data Panel to see which KBase Apps work with that data type.

For example, our annotated genome can be used in a number of analyses because it now includes the standard KBase annotations. If you click on the genome in the Data Panel, you see several apps that apply. Among them are the [*Insert Genomes into Species Tree* App](https://kbase.us/applist/apps/SpeciesTreeBuilder/insert_set_of_genomes_into_species_tree/release) that constructs a phylogenetic tree allowing you to see species closely related to your genome. You also can use the [*Build Metabolic Model* App](https://kbase.us/applist/apps/fba_tools/build_metabolic_model/release) to draft a metabolic model from the annotated genome.

## **Re-run Apps with the same or different parameters**

App input cells that have already been run have a “Reset” button that allows you to redo the analysis, with the option to change any of the parameter settings (including, perhaps, the input data) before rerunning. Rerunning an App will overwrite the information in the App cells, but will not overwrite any data objects in your Narrative *unless* you use the same output object name again.

## **Report errors**

If a job fails, you will see an error message in the output cell and/or in the *Result* tab. If you need help figuring out what went wrong, please [contact us](/troubleshooting/report) with details about what you were doing when the error occurred and what the error message said. Including the contents of the *Job Status* tab and the URL of your Narrative will help us debug the problem. We are here to help and appreciate your feedback and error reports!


# Job Browser

The Job Browser is one of the buttons on the left sidebar menu in KBase. In addition to viewing job status within a Narrative, the Job Browser allows you to monitor and manage your jobs across Narratives. By default, it shows all jobs submitted within the Previous Week. You can change the timeframe to the Previous hour, 48 hours, month, year, All Time, or a custom range of dates using the Time Range drop down menu.

<figure><img src="/files/WEMNayj1dcGfdMRPtAMH" alt=""><figcaption></figcaption></figure>

Filter jobs through toggling filters. The pop-up will  allow you to filter the jobs by *Created*, *Queued*, *Running*, *Completed*, *Error,* and *Canceled.* You can check as many filters you want and apply other terms. Applying these filters will update the page. Otherwise, the table does not auto-refresh.

<figure><img src="/files/3tYfNq6Z4mVDs4HGGfRu" alt=""><figcaption></figcaption></figure>

The results display basic information for the jobs such as the Narrative they are located, App, date Submitted, time Queued, Run time, current job Status, Server, and if the job was Canceled. The name of the Narrative is linked, and clicking on it will open the Narrative in a new tab. The App ID is linked and clicking on it will open a new tab with the App Catalog page.

### Job Status

The overall Job status is stated in the Status column. To inspect the Job Log, click the <img src="/files/DIq3MORpJ9x1RwUWc9vH" alt="" data-size="line"> icon to open the table and show the contents of the log. Click the 'x' or 'Close' buttons to close the table. Canceling the job will prompt you to make sure you didn’t click it accidentally. Jobs that have been canceled may continue to show up as “Queued” or “Running” until they clear the system.

### Job Log

The Job ID and worker node are easily located and the log can be scrolled through within the pop-out. Job Logs can be downloaded in CSV, TSV, JSON, and TEXT formats using the download (downward-facing arrow into tray) icon.&#x20;

<figure><img src="/files/RJpj0XNSfjlZb5cpesyi" alt=""><figcaption></figcaption></figure>


# Revise Your Narrative

You can browse a list of your Narratives in the *Narratives* tab:

<figure><img src="/files/py3Vt5kXSWnvpuT0fsHJ" alt="" width="375"><figcaption><p>View Narratives, create a new Narrative workspace, or copy a Narrative.</p></figcaption></figure>

Your Narratives are shown in the *Narratives* tab in reverse chronological order, with the most recently updated Narrative first. After your personal Narratives, more Narratives owned by others and shared with you are listed.

## View Narrative history and revert to earlier versions

KBase stores a history of everything you’ve done to your Narrative, so you are able to revert changes and go back to an earlier version. When you hover your cursor over one of your Narratives in the Narratives panel, three icons will appear.

<figure><img src="/files/-LvRfRhlAM9Mo2--x6lZ" alt="" width="375"><figcaption></figcaption></figure>

The first of these icons (counterclockwise encircling arrow) lets you View Narrative history to revert changes or access prior saved versions of the Narrative. Here you can choose an earlier version of a Narrative to revert to from the saved versions.&#x20;

## Add text to your Narrative

Adding text cells to your Narrative to explain what you were doing and what the results might mean will help others understand your computational experiments, as well as helping to remind you what you were doing. You can add formatted text to your Narrative in “Markdown Cells”. Markdown is a markup language (one example is HTML). It’s pretty easy to learn but does require knowing some syntax. A good [Markdown CheatSheet](https://github.com/adam-p/markdown-here/wiki/Markdown-Cheatsheet) can be found on GitHub.

You can add Markdown cells to a Narrative using the translucent blue buttons at the bottom right of the main Narrative panel.

<img src="/files/-M6BW8l96ScvbMKw3ydW" alt="" width="188">

The button on the right is the “Add Markdown Cell” button and will add a Markdown cell under the currently selected Narrative cell.

New Markdown cells contain the placeholder text “\[Empty Markdown / LaTeX cell]” (LaTeX is a formatting language). To edit a Markdown cell, double-click it to enter edit mode. When you’re done, press Shift-Return to exit the cell.

The ">\_" button will Add a Code Cell, which is beyond the scope of this user guide. TIP: To use a code cell press Shift-Return within the cell to execute the code.

## Cell commands: move, delete, etc.

Via the arrows and menu items in a Narrative cell, you can move the cell up or down, collapse or expand it, show the code that runs the analysis, get more info about the app, or delete the app cell. The cell menu can be opened by clicking the “…” icon in the upper right corner of a cell.<br>

<img src="/files/-M6qELqIM2z-dQE-PK9F" alt="" width="188">

## Rename or delete a Narrative

To rename a Narrative, click on its name at the top of the Narrative window. A window will pop up to let you change the name.

If you decide to delete a Narrative that you own, click the trashcan icon. Delete with caution, because currently there is no way to get your deleted Narrative back.

<figure><img src="/files/ZRALqJ4OmHY0j8YKlEk4" alt="" width="375"><figcaption><p>Delete a Narrative from the My Narratives Tab</p></figcaption></figure>

## **Save your work!**

Although your Narrative is periodically autosaved, it is always a good idea to save your work frequently, using the Save button at the top right of the screen.

<img src="/files/-M6qETpvBJ12hUN2quIb" alt="" width="375">

Once you are happy with your Narrative, you will probably want to share it with others. The section [Share Narratives](/getting-started/narrative/share) describes how to do that.


# Format Markdown Cells

To edit the appearance in Markdown cells for publication, teaching, and to add navigation within the Narrative, use a combination of Markdown and HTML.

Add formatted text to your Narrative using [Markdown cells](/getting-started/narrative/revise#add-text-to-your-narrative) to explain your process, interpret results, and communicate to others for sharing and publication and more.&#x20;

Markdown cells use markup languages of Markdown and HTML. Both are pretty easy to learn but does require knowing some syntax. Not every formatting option will work in the Narrative Interface.&#x20;

### [Narrative Formatting Cheat Sheet](https://narrative.kbase.us/narrative/69716) to copy and save

### Resources for Markdown syntax

* [Markdown Cheat Sheet](https://www.markdownguide.org/cheat-sheet/)
* [Markdown Guide](https://www.markdownguide.org/basic-syntax)

### Resources for HTML syntax

* [HTML Cheat Sheet](https://htmlcheatsheet.com/)
* [HTML Guide](http://www.simplehtmlguide.com/essential.php)
* [HTML Color Codes](https://htmlcolorcodes.com/color-picker/)


# Share Narratives

How to share your KBase Narrative workspaces with individuals and Orgs.

Narratives are *collaborative* (multiple people on your project can work on a Narrative together), *shareable* (you can share your Narrative with specific people or with everyone), *publishable* (a finished Narrative is similar to a research paper, and can be cited as a URL), and *reproducible* (you can repeat or even change another scientist’s computational experiment if they publish it as a Narrative). We believe that you will ultimately help both yourself and other researchers by sharing your Narratives so that others can see what you were thinking, what you did, and what you concluded from your analyses. However, **all of your Narratives are&#x20;*****private (viewable only to you)*****&#x20;until you decide to share them.**

{% embed url="<https://youtu.be/T07ugYD8UpU>" %}
A quick 4-min Tutorial on how to share Narratives.&#x20;
{% endembed %}

## How to Share a Narrative

You can share a Narrative that you own with any set of people you choose, or with all KBase Users. To share a Narrative, click the “share” button near the top right and select the "Manage Sharing" option in the menu.&#x20;

<img src="/files/-M6qA2Bivl59SmceNGIA" alt="" width="375">

You will now see the Change Share Settings pop-up window. Here, you can see the people you have already shared with (right now, only you).

<img src="/files/-M6BtAZLP4pNCdzLskG3" alt="" width="563">

If you want to share your Narrative with everyone, you can click the “make public” link near the top of the sharing panel.&#x20;

<figure><img src="/files/18tLl6jSh2dBiv1hC10p" alt=""><figcaption></figcaption></figure>

### Users tab

If you only want to share it with a few people, you can type their name or username into the “share with” box. Once you have typed something that matches a name in the KBase User Database, the person’s full name will appear. Click the down arrow to the right of the “share with” box to assign permissions to that user–you can allow them only to view your Narrative, to edit it, or to edit and share it.

For example, suppose you want to share your Narrative with kbasehelp so that KBase staff can help you troubleshoot it. You can start typing any part of the username or real name and a list of users that match will appear:

![](/files/-LvRgSQfKUQqHlO8S8yh)

The default permission level for sharing is “View only”, which lets the user you’re sharing with see the Narrative but not let them edit it or share it with others. If you want to share your Narrative with a collaborator and give them more privileges, click the down-arrow to see the permission levels and possibly choose a different one.

<img src="/files/-M6qGpqJGPPOteHZvq7V" alt="" width="563">

You can add multiple users to share with. Click the “Apply” button to add the chosen user(s) to the sharing list for your Narrative. The usernames and permission levels for the users you shared with will appear at the bottom of the Share panel.

### Orgs tab

To share your Narrative with a specific Organization, select the Orgs tab. Here you can add a single or multiple Organizations to share with a group of users.&#x20;

<img src="/files/-M6qH_Hp04tsV7bdY-dy" alt="" width="563">

Click the “Apply” button to add the chosen Organization(s) to the sharing list for your Narrative. The Organizations and permission levels for the users you shared with will appear at the bottom of the Share panel.

When you’re done sharing, click the "x" icon to close the Share panel.

## Editing access permissions

You can change the access permissions at any time by clicking the Share button again. The up/down arrows next to each user’s permissions let you change the permissions.

Note: Unless you alert another user that you have shared a Narrative with them, they will not notice a newly shared Narrative until they see your shared Narrative in their “Narratives that have been shared with you” tab or click on your name in her collaborator list.

If you are working on a shared Narrative, and a collaborator is working on it at the same time, watch for a red box at the top of the screen that says “Narrative updated in another session.” You can update to the latest version by refreshing the page, but note that any changes that you made to the Narrative will be lost, and if you try to save your changes, you might overwrite your collaborator’s changes.

## Publishing a Static Narrative

You can share you Narrative in publication ready, uneditable format that is accessible outside of KBase by creating a static Narrative. Once you are ready to share the Narrative, there are two steps.

&#x20;1\. Click on the share icon in the menu and click on Manage Sharing in the dropdown menu. In the pop-up window click 'make public?' to change the privacy settings. Click the 'x' icon to close the pop-up window.&#x20;

![Gif of opening the Manage Sharing pop-up window](/files/-MYC0PnEAVtcIKc4VA5d)

2\. Return and click on the 'share' icon, go to 'Manage Static Narratives in the dropdown menu. The 'Static Narratives' pop-up window will open and show the status if any versions of this Narrative have been made static. To create a static Narrative of the current version, click on the blue 'Create static narrative' button.&#x20;

![Gif of creating a Static Narrative](/files/-MYC0UxXKpzopkE9KIym)

3\. Access the static Narrative and share the link by clicking on the 'Existing static Narrative' link within the pop-up window.&#x20;

<img src="/files/-MIUl0cPhqAFAiVv7FTy" alt="Image of Static Narratives pop-up window." width="563">

{% hint style="warning" %}
If the Narrative is not public, you will not be able to generate a Static Narrative.&#x20;
{% endhint %}

<figure><img src="/files/TILSb7vk9Jqp5cQiAQJF" alt="" width="563"><figcaption></figcaption></figure>

{% hint style="info" %}
With a static Narrative, you can request a DOI to include and cite your Narrative in a publication. Request a DOI by emailing us at <engage@kbase.us>.&#x20;
{% endhint %}


# Linking Static Narratives to ORCID

How to add KBase static Narrative DOIs to your ORCID account.

Adapted from the article [DOE OSTI Search & Link Wizard Added to ORCID by  Stephanie Gerics and Carly Robinson](https://info.orcid.org/doe-osti-search-link-wizard/)

## Add Works using the DOE / OSTI Search & Link Wizard

1. Sign in to ORCID
2. Scroll down to the Works section of your ORCID record and hover over the '+Add works' in the menu and select 'Search & link'

<img src="/files/-MdUJ13_W1lXdgYkNkF4" alt="" width="563">

3\. Under 'Link Works', select DOE / OSTI from the list of member organizations.&#x20;

<img src="/files/-MdUJkG7DQEZxKZK8wz5" alt="" width="563">

4\. While using the DOE / OSTI Search & link wizard for the first time, you will need to Authorize access to approve DOE / OSTI to make changes to your ORCID Works.&#x20;

Once DOE / OSTI is authorized, the OSTI.GOV page will appear with the green ORCID iD icon next to the user’s ORCID iD name, as shown below. The ORCID Account Integration page will display *Authorization successful!*, the user’s *ORCID Account Details,* *ORCID Works in OSTI.GOV,* and directions for adding works to the user’s ORCID record using OSTI.GOV’s search. Browse DOE / OSTI search results to identify records authored, coauthored, or contributed by the user

<img src="/files/-MdUK1lb90t0mzl1iu1Q" alt="" width="563">

5\. Once you find records in the search results to add to your ORCID record, select 'Add to ORCID Works.'

6\. A pop-up text box will appear and confirm that the user wants to add this record to ORCID Works. Select 'Add to my ORCID Work&#x73;*'*.&#x20;

7\. In the search results, works added by the user to their ORCID record will now display 'In your ORCID Works' at the bottom right of the popup.

8\. To view your list of works added to your ORCID record, navigate back to the 'ORCID Account Integration' page. Select the user name next to the green ORCID iD icon at the top of the page.&#x20;

9\. Click on the ORCID Works in OSTI.GOV to see the works integrated with your account.&#x20;

<img src="/files/-MdULVW8uBRF6IqygyA2" alt="" width="563">


# Access and Copy Narratives

## View Narratives created by or shared with you

The Narratives tab has two sections that list Narratives that you have created (*My Narratives*) and, below that list, those that have been shared with you by other users. The *Shared With Me* list includes only those Narratives that were explicitly shared with you, not those that are viewable by everyone (public Narratives). The public Narratives can be accessed from the Narratives dash.

If you click on one of the Narratives in your Narratives panel, it will open in a new web browser window. If it’s a Narrative that you own, you will be able to run it or edit it. However, if it’s one that someone else owns and shared with you without write permission, you won’t be able to edit or run it–it will open in *View-only mode*.

### **View-only mode**

The key difference you will notice when a Narrative opens in View-only mode is that there is no Analyze panel on the left side–you won’t see any data objects or apps. This is because you can’t modify or run a Narrative that was shared with you with only View permission.

<figure><img src="/files/FSL9LfOllQeT40Kab3MB" alt=""><figcaption></figcaption></figure>

Why can’t you run a Narrative owned by someone else? Because, as we’ve seen, running an app usually creates new data objects in your Narrative, which requires write permission. If you want to run a Narrative owned by someone else who has not given you write permission, you will need to copy it to your account.

The blue box with a small white arrow on the top left of a View-only Narrative opens the Data Panel (so you can see the data objects in this Narrative) and the Narratives list (so that you can go to another Narrative).

{% hint style="info" %}
You can switch a Narrative that you own to view-only mode by clicking the pencil icon under the title of the Narrative.
{% endhint %}

If you ever want to print out a Narrative, switch it to view-only mode first, so it doesn’t waste the left side of the page with empty space.

## Copy a Narrative

To copy a Narrative that has been shared with you, so that you can run or change it, you can use the blue “Copy This Narrative” button at the top of the Narratives panel (or, in View-only mode, at the top right of the screen). This will copy the Narrative that you are currently viewing (which is marked with a green dot in the list of Narratives). You can choose a name for the copy (or leave the default name, which is the original name with “Copy” at the end).

<figure><img src="/files/JsPucUcYptq7T81cQrv0" alt="" width="563"><figcaption></figcaption></figure>

When you click Copy, the Narrative and its associated data objects are copied to your account, and you should see the new copy appear at the top of “My Narratives”.

#### **Copy of a copy warning**

Please be aware that if you copy a copy of a Narrative, some things may not work right in the copy of a copy. We are working on resolving this problem.

You can copy any Narrative by hovering your cursor over its name in your Narratives panel to reveal the icons and clicking the Copy icon (the one that looks like two sheets of paper).

<img src="/files/T13dbN7Q5MRsxuE7IdA0" alt="" width="375">

The copy is now yours–you can open it, run it, edit it, or even share it with others (if the owner of the original Narrative gave you sharing privileges). After you’ve copied a Narrative, you can run all the steps to reproduce the computational experiment, or try changing the parameters or input data to alter or perhaps even improve the results.

Copying Narratives is the key to leveraging KBase’s commitment to reproducibility and reusability. By making it easy to build on previous work and go through rapid cycles of analysis, KBase aims to accelerate the pace of systems biology research.


# Organizations

How to access, search, and create KBase Organizations for labs, collaborating researchers, courses, and more.

## What is a KBase Org?&#x20;

In science, people are members of labs, teams, groups, collaborations, and projects that work together with shared data and analyses. In KBase, organizations are a way for teams of scientists to share their data and associated analyses that are in the Narratives they create with each other as a group. Organization members can see information about the team and a list of the Narratives associated with the organization to which can request access. View-only access is granted upon request while all other access to the Narrative is granted by the Narrative owner. KBase users can be members of more than one organization and Narratives with the associated data might also be added to more than one organization.&#x20;

{% embed url="<https://youtu.be/fZ7CMggm1qM>" %}
7 minute video tutorial on KBase Organizations
{% endembed %}

## Creating an Organization

On the Orgs page, click "Create Organization" button on the right side. This will open the Create Your Organization page. Input the information and Save.&#x20;

![](/files/-MS_RsIiqo8gAq4fLN8_)

### To create an Organization, you will need:&#x20;

* Name ⏤ the displayed Organization Name
* ID ⏤ the unique URL for the Organization; can only contain lower case letters (a-z), numeric digits (0-9) and the dash "-"
* Logo URL (optional) ⏤ include the link to a publicly available image&#x20;
* Home Page URL (optional) ⏤ input your lab or group's website URL to link and share
* Hidden? ⏤ check for the Organization to be *Hidden - will be visible **only** for members of this organization* or leave unchecked to remain *Visible - will be visible to all KBase users*.
* Research Interests ⏤ list your Organization's research interests here
* Description ⏤ provide a description of your Organization

### Editing an Organization

<img src="/files/-MS_Pykk9vMpMSX_Dy3R" alt="" width="563">

There are options to edit and share an organization after it has been created.&#x20;

## Joining an Organization

* If an invitation appears in your feeds, click on the ID of the organization. A window will pop up with options to Accept or Reject the invitation. Click on your preference.
* Without invitation, search for the organization and click on its name. Click on the blue button (upper right) to ‘Join Organization’. You cancel this request if needed. A request to join will be sent to the owner and group administrators. An acceptance or rejection will be added to your Feeds. Note: Private organizations will not appear.&#x20;
* If you know the organization URL (starts with <https://Narrative.kbase.us/#org/>) go to the webpage and request membership.

## Viewing Narratives in an Org

* If you haven't opened the Narrative before, there will be a "Click for Access/View Access" button
  * Click the "Click for View Access" button
  * You will be able to open and view the Narrative when the "No Access" line disappears
  * To open a Narrative, click on the Narrative name

<img src="/files/TkSA7ihOgCQWJNJUSqiM" alt="Access a Narrative by clicking the &#x22;Click for View Access&#x22; button to be able to view the Narrative. " width="563">

* There are several icons that may appear in the Narrative cell:
  * **globe** <img src="/files/-MFqXVPBP09wQOkw0Mxr" alt="" data-size="line"> - Narrative is shared publicly with all KBase users
  * **pencil** <img src="/files/-MFqXf2NI_nw5SYZoOKh" alt="" data-size="line"> - Narrative is shared with you with *edit access*
  * **open lock** <img src="/files/-MFqXm9ZLqlTbFbuSnKQ" alt="" data-size="line"> - Narrative is shared with you with *access* to edit and share
  * **closed lock** <img src="/files/-MFqXqRl-6EIAJ2KUufK" alt="" data-size="line"> - private
  * **eye** <img src="/files/-MFqXuMFO4CW5YPv3eDx" alt="" data-size="line"> - View-Only access
  * **crown** <img src="/files/-MFqXwvzHAm2CY1yOn66" alt="" data-size="line"> - you are the owner

## Administering an Organization

* Add Narratives by either the "+ Add a Narrative" button or [share Narratives](/getting-started/narrative/share#users-tab) with an Org.&#x20;
* Invite members by clicking the invite button on the right above the list of members.

  In the Search box ‘Select User to Invite’, type in the user’s name or KBase ID. A list of users will appear below the search box. Click on the user to invite and the name will appear on the right with more information that can be used to confirm it is the correct user. If correct, click on the blue ‘Send Invitation’ button.&#x20;
* On the right side of the organization page is a list of organization members. Above the list of members is a tab called ‘Requests’. There is an "Inbox" for a list of Users requesting to join and an "Outbox" of owner/administrator invitations for pending invitations for users to join that the owner or administrator made. Requests can be accepted or rejected by clicking the respective buttons. Invitations can be cancelled at any time, but will also disappear after a length of time.&#x20;
* Owners and administrators can can promote members to administrator. In the membership list, click on the ‘...’ to the right of the member’s name and click on ‘Promote to Admin’.
* Members can be removed by click on the ‘...’ to the right of the member’s name in the membership list and click on ‘Remove Member’. The user is not notified that they have been removed.
* &#x20;There is no mechanism to delete an organization. You can un-associate all Narratives, remove all members and set the group to “Private.” \ <br>


# FAQs

Frequently Asked Questions about KBase.

## Is KBase free to use?&#x20;

Yes! KBase is completely free to use and is supported by the Department of Energy.

## I'm getting an error. What do I do?

You can check [the job log](/troubleshooting/job-errors/common/job-log) and refer to our [troubleshooting guide](/troubleshooting). If you do not see a reason for the error there, please make a ticket on our [Help Board](https://kbase-jira.atlassian.net/).

## Is it possible to obtain publication-quality figures by using these tools?

Some of the Apps can produce publication-quality images that can be views by opening in a separate window. We hope to add more in the future.

## I have too much data to move.&#x20;

Is there a better way for handling large datasets? For large files or collections of files, we recommend using Globus. We also have a [guide for importing using Globus](/data/globus).

## Can I import more than one file into a Narrative at the same time?

Yes, bulk import allows you to check multiple files in the staging area and import them all by clicking "Import Selected." More instructions on this feature are [described in the Narrative Guide](/getting-started/narrative/add-data#using-bulk-import).&#x20;

## How many users can view a Narrative?&#x20;

There is no limit to the number of users that can be invited to view or edit Narratives. However, it is possible for users who are accessing the Narrative simultaneously to overwrite each other's changes. For that reason, we recommend only one person edit the Narrative at a time.

## I work with datasets with many samples, outputs from each step often mixed together and often hard for me to select the ones for the next step. How do I keep them organized?&#x20;

The best way to organize many inputs is to create sets. There are Apps for building these sets, such as [Build AssemblySet](https://narrative.kbase.us/#catalog/apps/kb_SetUtilities/KButil_Build_AssemblySet/release) that allow you to group those outputs.

## What languages are supported in code cells?

Currently we only support code cells in Python.

## How does KBase compare with the standalone tools available with command line interface? Does KBase always use the most recent version?

Provided the same data, parameters, and tool version are consistent, the results are the same. We make every attempt to use the latest versions of Apps and databases, but they do not always correspond to the most recent version.

## How do I access advanced parameters?&#x20;

You can view some advanced parameters depending on the app by clicking the Advanced Parameters button. For some apps there are few or no advanced parameters. In creating these Apps we aim to balance configurability and usability. As a result, for some Apps we do not make certain rarely changed parameters visible.

## How do I get a tool I like added?&#x20;

There are a couple venues:&#x20;

1\. Submit a Help Desk New Feature Request. We can't promise if or when new feature requests are added, but user input helps us plan our future development.&#x20;

2\. Add as a 3rd-party developer: If you are a developer of an app, you can find more information on [the developer's page](/development) and you can read the documentation for our [Software Development Kit](https://kbase.github.io/kb_sdk_docs/) . If you are not the developer, you can reach out to the developer and ask them about wrapping it for KBase.&#x20;

Alternatively, a beta version of the App might already be in KBase. In the Apps panel, click on the "R" button to switch from release to beta. When a "B" appears in place of the "R", you are viewing beta Apps.

## What pipelines does KBase provide?&#x20;

&#x20;KBase does not provide *pipelines* per se; we try to be as flexible as possible. We do have [example Narratives](/workflows) that replicate the paths many users follow, but we encourage everyone to plot their own course.

## Will my data be made public?

All user data in KBase is [private by default until you share it](https://www.kbase.us/data-policy-and-sources/). You can keep all your data private if you wish. We do encourage users to make data public in the interest of FAIR - [Findable, Accessible, Interoperable, and Reusable](https://www.go-fair.org/fair-principles/) - data practices.

## What organisms are available in KBase?&#x20;

In KBase as a whole, anything can be uploaded, the only exception being that human data is not permitted. There are connections to some bacterial, plant, and fungal species in the Public Data tab, but any user data or data found from public sources other than those integrated into KBase can be uploaded. On an app-by-app basis, some tools may work sub-optimally or not at all with certain species.

## Is there an API to allow access to the methods from the command line?&#x20;

Command line is not supported for the narratives. However, you may find more flexibility with the KBase Software Development Kit (SDK).

## Can I get an in-person KBase workshop at my institution?&#x20;

If you would like an in-person workshop, contact us by email at <engage@kbase.us>.

## Can I see the background code for KBase?&#x20;

Yes, the relevant Github repo for each App can be found by following the links in the App Catalog. You can access the version and parameters that were used during an App run by clicking on the ellipsis button and clicking "Show code."

## **What is the difference between Public and Private organizations?**

* The existence and contents of Private organizations are only visible to members of the organization. The organization will not be findable by non-members.
* Any KBase user can see a list of all Public organizations plus a description of the organization. Non-members are given the option to request membership. Members of the organization can access more information than non-members.

## **Why can't I open a Narrative in an Org?**

Click on ‘Click for Access’ and you are automatically granted View access to the Narrative. You can then click on the linked Narrative to open.&#x20;

## What if I lose access to my account?&#x20;

You can log into KBase through Google, Globus, and ORCiD. It is recommended to always have at least one personal account linked to your KBase account, so that in the event you lose access to an identity supplier i.e., Google, Globus, or ORCiD because of an email provider change, then you will still have another method to login.&#x20;

If you do lose access completely, please submit a ticket through the [Help Board](/troubleshooting/support#contact-us).&#x20;

<br>


# Manage Your Account

## Create or modify your user profile

Every KBase User has a user profile. Filling in your KBase user profile helps you become part of the social web that KBase is building, making it easier to find and collaborate with other scientists and share data and Narratives. You are not required to complete your user profile in order to use KBase, but we encourage everyone to do so.

To see and edit your user profile, go to the Menu sidebar on the left and click the “Account” button on the left side.

<figure><img src="/files/AcCAr5yAgDVWD7gTi2HC" alt="" width="90"><figcaption></figcaption></figure>

Your profile page will open, looking something like the screenshot below. You will be able to enter or change information including your organization, department, job title, research interests, avatar, and more. Required fields are marked with a red "\*." You must fill in all required fields in order to save your profile (using the Save button on the upper right of the page).

<figure><img src="/files/FAQ2NR1Lbjj6YT8kmGaF" alt="" width="563"><figcaption></figcaption></figure>

Below the "Save" button is a small preview showing what your profile page looks like to others. You can click the “Open Your Profile Page” link to see a bigger version, along with a list of your Narratives and collaborators (other users with whom you have shared Narratives, or who have shared Narratives with you) and a search box for searching for other users by name or username. As you start typing in the box, you will see a list of users who have that text in their first or last name or their username. Click one of the matches to see that user’s profile.

The **Account** tab lets you see basic information about your KBase account. The only field that can be edited there is your preferred email address. Note that other users cannot see your email address; it will be used only by KBase staff to contact you occasionally with important information.

The [**Linked Sign-In Accounts**](/manage-account/link-accounts) tab lets you see and [manage](/manage-account/link-accounts) which external (Globus or Google or [ORCiD](/manage-account/link-orcid)) accounts are linked to your KBase account.

The **User Profile** editor allows you to edit your position, organization, and location. You can also input your Research Interests, Research or Personal Statement. It also shows which KBase *Organizations* you belong to and any Affiliations.&#x20;


# Linking Accounts

Linking Accounts can be helpful for data transfers, ease of signing in, and sharing data. You can link your KBase Account to Google, Globus, and ORCiD.&#x20;

To link an account, first sign into KBase.

On the left sidebar menu, click the “Account” icon, and then choose the “Linked Sign-In Accounts” tab.

<figure><img src="/files/Y5ywWduKFyTjTcLQeULC" alt="" width="375"><figcaption></figcaption></figure>

At the bottom of the page, choose an account type from the pulldown list of identity providers, and then click the “Link” button.

<figure><img src="/files/VbxVBHZ0bkWH8267Snpz" alt="" width="375"><figcaption></figcaption></figure>

The next screen will prompt you to sign-in using your to the chosen account type to link.

After you select the "Link" button, you will be returned to the Linked Sign-in Accounts screen, where you will see your newly linked account.&#x20;

### Example

If you have a Google (gmail) account, you can link it to your KBase account.

First sign in to KBase with your current login method. If you are not already signed in to KBase, go to the [sign-in page](https://narrative.kbase.us/), click the Sign In button and sign-in as usual.

On the left sidebar of your menu, click the “Account” icon to access the Account Manager page and then choose the tab labeled “Linked Sign-In Accounts”.&#x20;

<figure><img src="/files/OjDgroaDkIEcLJYWsuan" alt="" width="375"><figcaption></figcaption></figure>

At the bottom of the tab, choose “Google” from the pulldown list of identity providers.

<figure><img src="/files/j5wSA22E2u6yDLX6RUhH" alt="" width="375"><figcaption></figcaption></figure>

Click the “Link” button. The next screen will prompt you to choose a Google account to link (some people have more than one).

<img src="/files/-M5xwmzrbAEWmNyf8zOx" alt="" width="375">

You are now ready to link your Google account to your KBase account!

<img src="/files/-LvNq2FO_UVROYiigrmF" alt="" width="563">

After clicking the "Link" button, you will be returned to the Linked Sign-in Accounts screen, where you will see your newly linked Google account.

### The next time you sign in with your Google credentials

Now when you want to sign in to KBase, you can also use the “Sign in with Google” button.

<figure><img src="/files/20aOR74ri10bKF5IPuRC" alt="" width="563"><figcaption></figcaption></figure>

This will bring you straight to the “Choose an account” Google page:

<div align="center"><img src="/files/-M5xwUA7NDFy5_kf24qB" alt="" width="375"></div>

After you choose your linked Google account, you’ll be signed in to KBase!&#x20;

{% hint style="info" %}
If you had signed out of this particular Google account, you will be directed to sign into it again before you login to KBase.
{% endhint %}

We hope you appreciate the convenience of being able to sign-in to KBase with your Google credentials. If you encounter any difficulties or have any questions, please feel free to [contact us](https://www.kbase.us/support/)!


# Sign into KBase with ORCiD

If you have an[ ORCiD](https://orcid.org/) account, you can sign into KBase using your [ORCiD <img src="/files/-M9K9Aepzk1FNZ3jYzOv" alt="" data-size="line">](https://orcid.org/) and link it to your KBase account.

{% hint style="info" %}
KBase is an [ORCiD Member Organization. ](https://orcid.org/members/0016f00002ZLyhNAAT-kbase)
{% endhint %}

If you are already a KBase User, first sign into KBase through your Globus or Google accounts. If you are not already signed in to KBase, go to the [sign-in page](https://narrative.kbase.us/), click the Sign In button, and click either “Sign in with Google”  or “Sign in with Globus” buttons and go through the normal sign-in process until you get to Narratives.

On the left sidebar menu, click the “Account” icon, and then choose the tab labeled “Linked Sign-In Accounts”.

<figure><img src="/files/Y5ywWduKFyTjTcLQeULC" alt="" width="375"><figcaption></figcaption></figure>

At the bottom of the tab, choose "[<img src="/files/-M9K9Aepzk1FNZ3jYzOv" alt="" data-size="line">](https://orcid.org/) ORCiD” from the pulldown list of identity providers, and then click the “Link” button.

<figure><img src="/files/lSsPLy6JLD3bGNOtXHOK" alt="" width="563"><figcaption></figcaption></figure>

The next screen will prompt you to Sign in using your ORCiD account to link.

<img src="/files/-M5y-xxLZ1lIhMjUTqO5" alt="" width="375">

After you click that Link button, you will be returned to the Linked Sign-in Accounts screen in your Account Manager, where you will see your newly linked ORCiD account.

You can now sign-in to KBase using your ORCiD ID and [integrate DOIs](https://info.orcid.org/doe-osti-search-link-wizard/).&#x20;

![](/files/-M6q8LyRhzGnjP2buTsa)

We hope you appreciate the convenience of being able to sign in to KBase with your ORCiD credentials. If you encounter any difficulties or have any questions, please feel free to [contact us](https://www.kbase.us/support/)!


# Link ORCiD Records with KBase

The KBase team is developing new features to provide author credit for data and analysis on the platform. Many users already use their ORCiD account to conveniently log into KBase, but this is an additional step to enable some new credit features. We’re encouraging all KBase users to link their ORCiD records to their account to better support author credit. This is necessary even if you already use ORCiD to log in.  This will allow for creating publication records in your [ORCiD profile when you publish a static Narrative](https://docs.kbase.us/getting-started/narrative/link-doi) with a DOI in KBase. In the future, this will enable you to automatically populate information into your KBase account from your ORCiD Profile as well. You can link your ORCiD Records to your KBase account in a few simple steps:

1. Log into your KBase account
2. Navigate to Accounts, then to the [ORCiD Record Link](https://narrative.kbase.us/account/orcidlink) tab&#x20;

<figure><img src="https://lh7-rt.googleusercontent.com/docsz/AD_4nXdvrO7wwfWgw0gC94a9pOEmMM33DqEx6I6OgsVFHpLrMQgabruAx8ByQ6kBrAZ4THx2EcRInz_KvKI67Sq_7I4WOsOvT1c3t3TqKrYn2dpR6c908AQv67VLDdV3TMZpOwMniv8PpQ?key=WIHcTPuD8Hizcd6-KQ2_zw" alt=""><figcaption></figcaption></figure>

3. Click the “Create ORCiD Record Link” button
4. When prompted, sign into your ORCiD account, and click “Authorize Access”
5. Back in KBase, click “Confirm ORCiD Record Link”

   <figure><img src="https://lh7-rt.googleusercontent.com/docsz/AD_4nXfgEZRmgiVtjCukAaH_9rq-NQX3xAKFGcnjmtf9sArPyrVQz8N8q9g0Mn687NW15MYSnusi4kE40QsU92o4GnqZwRtGwJ16-yVfSs0D1V2akLtAM3x9WA31Z7t6LdDDVqMRCGi-BQ?key=WIHcTPuD8Hizcd6-KQ2_zw" alt=""><figcaption></figcaption></figure>
6. You should now see your active ORCiD Record Link. You can remove the link as well in necessary

<figure><img src="https://lh7-rt.googleusercontent.com/docsz/AD_4nXfzB63Q8VBCTk9wKpfEYwuHFOsVKxwfIjbbW2QEPjjo3QQZ7eiCTjHKDkOdCjXOh3KHVkIpJwpDEX8zI4sFlQFn-V8bLoDPe059a8JDkz22QpkZNosFzgwolfWjK8rKaAFFo2RU?key=WIHcTPuD8Hizcd6-KQ2_zw" alt=""><figcaption></figcaption></figure>

Now you have successfully linked your ORCiD Records to your KBase account! If you have any difficulties, please create a ticket on the [KBase Help Board](https://kbase-jira.atlassian.net/) and we will assist you.&#x20;

<br>


# Working with Data

Guides for working with data in KBase

## Contents

1. [Upload and Download Guide](/data/upload-download-guide)
2. [How to search, add and upload data in KBase video tutorial](/data/search-add-upload)
3. [How to filter, manage, and view data in KBase video tutorial](/data/filter-manage-view)
4. [Linking data and metadata with Samples](/data/samples)
5. [Using public data within KBase](/data/public)
6. [How to transfer data using Globus](/data/globus)


# Data Upload and Download Guide

This guide provides instructions for uploading different types of data to your KBase account and downloading output data from your analysis results.

### A Quick Video Tutorial on Uploading and Downloading Data

{% embed url="<https://youtu.be/pEKYZbVkyCc>" %}

### Contents

1. [Data Types](/data/upload-download-guide/data-types)
2. [Importing Data](/data/upload-download-guide/uploads)
3. [Assemblies](/data/upload-download-guide/assembly)
4. [Genomes](/data/upload-download-guide/genome)
5. [FASTQ/SRA Reads](/data/upload-download-guide/reads)
6. [FBA Models](/data/upload-download-guide/fba-model)
7. [Media](/data/upload-download-guide/media)
8. [Expression Matrix](/data/upload-download-guide/expression-matrix)
9. [Phenotype Set](/data/upload-download-guide/phenotype-set)
10. [Amplicon Matrix](/data/upload-download-guide/amplicon-matrix)
11. [Chemical Abundance Matrix](/data/upload-download-guide/chemical-abundance-matrix)
12. [SampleSet](/data/upload-download-guide/sampleset)
13. [Compressed Files](/data/upload-download-guide/compressed-files)
14. [Downloading Data](/data/upload-download-guide/downloads)


# Data Types

Data within KBase covers a wide range of data types relevant to systems biology research, including genomes and their annotations, metagenomes, expression, and protein-protein interactions, inferred models of organismal and community metabolism and gene regulation, even geographical information about populations.

Remember to follow the KBase [Data Policy and Sources](https://www.kbase.us/about/terms-and-conditions-v2/#data_policy) agreement.&#x20;

{% hint style="danger" %}
When working with data in KBase and using Apps, remember that choosing the same output name for data objects will overwrite existing data objects of the same type with that name.&#x20;
{% endhint %}

## Data Type Descriptions

KBase handles data as objects versus files for interoperability with KBase Apps. The following are data types and their adjacent file types:&#x20;

* **Assembly** ⏤ FASTA files; extensions **.fasta, .fna, .fa**, **.fas**
* **FASTQ/SRA Reads** ⏤ Interleaved, Non-interleaved, paired-end reads, single-end reads; extensions **.fastq, .fq, .sra**
* **Genome** ⏤ GenBank or GFF3 with a FASTA file; extensions **.genbank, .gb, .gbk,** or **.gbff, .gff,** and **.fasta, .fna, .fa, .fas**
* **Metagenome** ⏤ GFF Metagenome with a FASTA file; extensions **.gbff, .gff,** and **.fasta, .fna, .fa, .fas**
* **FBA Model** ⏤ SBML, Excel, or TSV; extensions **.sbml, .xml, .tsv, .xls, .xlsx**
* **Media** ⏤ Excel or TSV; **.xls, .xlsx, .tsv**
* **Expression Matrix** ⏤ Excel or TSV; .**xls, .xlsx, .tsv**
* **Phenotype Set** ⏤  Excel or TSV; **.xls, .xlsx, .tsv .tab**
* **Amplicon Matrix** ⏤ Excel and FASTA; **.xls, .xlsx,** and **.fasta, .fna, .fa, .fas**
* **Chemical Abundance Matrix** ⏤ Excel, CSV, or TSV; **.xls, .xlsx, .csv, .tsv**
* **SampleSet** ⏤ Excel or TSV; **.xls, .xlsx, .tsv**

The [Data Type Catalog](https://narrative.kbase.us/#catalog/datatypes) lists and describes the key data types representing different classifications of biological data within KBase.


# Importing Data

Using the Data Browser in the KBase Narrative Interface.

## Using the Data Browser

The *Import* tab lets you drag & drop data files from your computer into the Staging Area to upload data into your Narrative. Visit [this page](/getting-started/narrative/add-data#uploading-data-from-external-sources) of the Narrative User Guide for general instructions on how to use the Staging Area.

![Data Panel and Import to Staging Area](/files/zr5OlDeUOpwlTsUKTjKp)

The first step in uploading your data is to locate the **Data Panel** along the left side of the Narrative Interface window and click the “Add Data” button, the circular “+” icon, or the arrow at the upper right of the panel to access the slide-out **Data Browser**. The Data Browser has several tabs that allow you to access data within KBase and the *Import* tab for uploading your own data.&#x20;

To close the Import panel and return to the Narrative Interface, simply click the “Close” button on the bottom right of the import panel.&#x20;

If you have a small screen, you might not be able to see that button. Another way to close the Data Browser is to click the arrow icon on the top right of your Data Panel (the same one that opens the Data Browser slide-out).

![Opening and Closing Data Panel](/files/t3U6YK528Qy2rAvRLSKs)

## **Importing Data**

The import type selection dropdown can detect the data type of the file based on its extension *where possible*. Some types will not automatically select, such as .fasta or .fastq. In these cases more than one importer can use the FASTA file extension. However, suggested importers will be listed at the top of the drop down under “Suggested Types.”&#x20;

{% hint style="warning" %}
Learn how to import multiple files under the [Bulk Import Guide](#bulk-import-guide)
{% endhint %}

For example, files with the extension FASTQ can import as a single interleaved file, as a pair of forward and reverse non-interleaved files to create a PairedEndLibrary, or as a single non-interleaved file to create a SingleEndLibrary.&#x20;

<img src="/files/NAvTOQwmAZS7pltz42va" alt="Suggested Types for Importing Data into the Narrative" width="563">

To import a file, make sure you have an import type selected and the check box is active, then click "Import Selected."

## **Drag & Drop Limitations**

The drag & drop from your local computer works for many files. There is a size limit around 2GB, which is dependent on your computer and browser. For larger files, use the [Globus Online transfer](/data/globus) or a direct upload with a URL.

<img src="/files/Z0wByozJnuJs78Ro6I1a" alt="Using Globus to upload large data files to the Staging Area" width="563">

## **Data Privacy**

Any data you upload to KBase is private unless you choose to share it. You can share any of your Narratives (including their associated data) with one or more specific users, or make it publicly available to all KBase users. Please see the [Share Narratives](/getting-started/narrative/share) page for more information and how to share data and Narratives. KBase policies on data are described on the [Terms and Conditions](https://www.kbase.us/about/terms-and-conditions-v2/) and [Data Policy](https://www.kbase.us/about/terms-and-conditions-v2/#data_policy).

The next sections of this guide describe the specific steps involved in uploading the currently supported data types, and show examples for each type.

## Bulk Import Guide

{% embed url="<https://youtu.be/NRp3coo1ry0?si=ngFsjkgUEEvq1RxG>" %}
How to Builk Import Files
{% endembed %}

To import multiple files, ensure all the files you want to import have an import type selected and the check box is active, then click "Import Selected."&#x20;

{% hint style="info" %}
Bulk import currently supports *one parameter set per data type*. If you need to upload files of a given data type with different parameters, perform a separate set of bulk imports for each parameter set.
{% endhint %}

![](/files/-MgGo7TNCJiyX4IKFjv_)

You’ll see a new import cell created for bulk imports. This cell contains a tab for each of the data types you selected from the Staging Area, which you can view by clicking on the available types within the **Data type column** on the left. This cell is a bulk wrapper for existing import apps, so you can fill out the parameters the same as you would for a single import.&#x20;

![Bulk Import App in the Narrative](/files/hKgU6tcx7y0Zlv0NY4qg)

Once you have filled out all file paths and parameters, run the import. While the import is running you can see all the logs and status details in the **Job Status tab,** formatted for the bulk import.&#x20;

Expand each of the child jobs by clicking on the corresponding line to view details, then expand further to see individual logs. When the bulk import cell is set to run, each child job for the individual data object will have a status, including **Actions** to *cancel* or *retry* individual jobs.

![Running Bulk Import App](/files/eZrqkqFdC1FNFvhaCXnP)

Alternatively, queued and running jobs can be cancelled using the bulk action dropdown. Or the retry action can be used for multiple import jobs within the bulk import cell. &#x20;

![Bulk Action Dropdown](/files/kLupEjAxoYXQuwqJE9AA)

When the jobs are completed successfully, you’ll be able to see all the information on the data object, repackaged for viewability in bulk. The first new feature is a table of all successful jobs. Clicking on object names adds a viewer widget to the Narrative. All the reports are available under the **Reports** tree. The full list is collapsed by default to conserve memory. The list can be expanded to show each report, and each report can be viewed in a separate window.&#x20;

![Result tab of the Bulk Import Cell](/files/FBXiuJ4zpiuKK8gp4NIx)

### What data types does this apply to?

While this will be applicable to all data types and file extensions in the future, bulk import currently supports the following:

* Assembly - FASTA
* SRA Reads&#x20;
* FASTQ Reads Interleaved - FASTQ Interleaved reads
* FASTQ Reads Noninterleaved - FASTQ paired-end reads or FASTQ single-end reads library
* GenBank Genome - GenBank
* GFF Metagenome - FASTA and GFF3

{% hint style="warning" %}
When selected files are not supported data types for Bulk Import within the Staging Area, separate, single import app cells will open in the Narrative for each type and file.
{% endhint %}

### What are the limitations?

This is a new feature in KBase, and the current release should be considered a beta version with future development still to come. Some of the most notable limitations are listed below. This is not a comprehensive list, but does contain the known bugs, issues, and limitations that are the highest priority for future releases. Please report any new bugs you find to the Help Board.

View [common bugs and limitations](/data/upload-download-guide/uploads/bulk-limitations).&#x20;


# Bulk Import Limitations

This page contains a list of the most commonly observed bugs and known limitations to the bulk import process.

## Known Issues and Limitations

* Issue: Only a single set of parameters can be used for a given type in a given bulk import.&#x20;

  Workaround: Subset your imports and run separate imports for each subset of data with a single set of parameters. Note that this only applies to parameters *within an import type*. For example, you have a set of Illumina reads, PacBio reads, and assemblies, you can import one of the sets of reads and the assemblies in one import, and the other type of reads in a separate import.
* Issue: Staging Area allows uploading incomplete files, leading to data corruption.

  Workaround: This is an existing issue in the staging area that exists for single-file imports as well. Ensure files have fully uploaded and the file size in the Staging Area matches the size of the file from your machine. You can use MD5 checksums to verify the file uploaded correctly. You can view the MD5 inside the Staging Area by viewing the info for the file (*shown below*). Then, verify that the MD5 in the Staging Area matches the MD5 on your local machine using the following command depending on your operating system (replacing "assembly\_1.fasta" with your file name):

  * MacOS: `md5 assembly_1.fasta` &#x20;
  * Windows: `certutil -hashfile assembly_1.fasta md5`
  * Linux: `md5sum assembly_1.fasta`&#x20;

<img src="/files/-Mk8koitIh4PomLxzdni" alt="Example for viewing file size information in the Staging Area to ensure the file is completely uploaded" width="563">

* Issue: File paths are not autofilled.&#x20;

  Workaround: During import pairs of files that go together, such as pairs of non-interleaved FASTQ files or FASTA and GFF Metagenome files, will need to be manually assigned to a forward or reverse read.&#x20;

![Selecting import pairs](/files/fUxpjahSKBVMBwpCak1O)

* Scientific Name input disappears and becomes undefined within the input parameters when switching between tabs.&#x20;

  Workaround: This is a visual issue and the selected Scientific Name will still be included when running the import - even though it is not searchable in the *undefined state*.&#x20;

  To verify the selected name is correct when the parameter does switch to the undefined state, 1) search for and select a different scientific name, 2) click between Data type tabs, 3) return to Data type tab and search for and select the original, preferred scientific name, 4) then Run the import.&#x20;
* Issue: Job Status/Logs/Results do not appear or do not update.

  Workaround: There have been infrequent bugs found in which logs failed to appear, jobs reported as successful but no report could be generated, and similar issues. Switching between tabs or reloading the page has been found to fix these issues.


# Assembly

How to upload FASTA files into KBase.

{% hint style="info" %}
An assembly file is a single file containing one or more contiguous DNA sequences in FASTA format. It can be uploaded to KBase from your local computer (with file extension .fasta, .fna, .fa, or .fas) or directly from a publicly accessible FTP or HTTP URL.&#x20;
{% endhint %}

“Assembly” is the KBase data type for assembled, unannotated DNA sequence contigs or  predicted coding sequences. If you want to upload annotated sequences in GenBank or GFF format, please see the [Genome](/data/upload-download-guide/genome) page.

![](/files/peJhfFWdhF0IubSgAbNa)

## Importing an Assembly from your computer

Using a file on your computer, open the [*Import* tab within the **Data Browser**](/getting-started/narrative/add-data)**.** Then drag & drop the assembly file into the Staging Area box or select from your computer files.

Open the *Import As* pulldown menu to the right of the filename in your Staging Area and select “Assembly.”

<img src="/files/stnlolbUmGavF3yVTLVQ" alt="Import Fasta Assembly File" width="563">

Make sure the correct file type is selected and the checkbox is active, then click the “Import Selected” button. The data slide-out will close and an app called “Import FASTA File as Assembly from Staging Area” will be added to your Narrative.

![Import Assembly / FASTA file Import App Cell](/files/LieDMQnDwqyeu12aLPl4)

The name of the Assembly file is filled in, as is a suggested name for the Assembly data object that will be created by the import (you can change the Assembly object name). Adjust the minimum contig length if needed and then click the green "Run" button to start the import. When the import is finished, your Data Panel will update to show the new Assembly object, and a report will appear in the Import App.

## Upload an Assembly from other sources

You can upload data into your KBase Staging Area using [Globus](/data/globus), or by supplying a URL for a publicly accessible FTP location, Google Drive, Dropbox, or a direct HTTP link to import into the Narrative. Options for adding data to your Staging Area are described [here](/getting-started/narrative/add-data).

<img src="/files/KzHe3U0JdzLl1GIJjBoi" alt="Drag and Drop, Upload with Globus, Upload with URL options to upload and import data into KBase" width="563">

### Bulk Import

Assemblies can be imported as one of the supported bulk import types. You can select multiple assemblies simultaneously from the staging area to import them at once. See the bulk import section of [the guide to importing data into the Narrative.](https://docs.kbase.us/getting-started/narrative/add-data)


# Genome

Formatting and uploading annotated assemblies and GenBank or GFF and FASTA files.

In KBase, a **Genome object is the annotated version of an** [**Assembly**](/data/upload-download-guide/assembly) or **annotated predicted coding sequences** and can encompass several types of feature calls. If you want to upload solely the DNA sequence from a FASTA file (without annotations), go to the [Assembly](/data/upload-download-guide/assembly) page.

{% hint style="info" %}
The Genome importer supports only GenBank and GFF formatted files.&#x20;

A GenBank-formatted input file should include sequence contig(s), feature calls (annotations), and taxonomy information for the organism. KBase parses the input file into two data objects: an assembly object with the sequence and a genome object containing the original feature calls and annotations.

GenBank-formatted files with no features can be uploaded as Genomes.

A *GenBank-formatted* file can be uploaded into the Staging Area from your local computer (files with.gb, .gbff, or .gbk extensions) or directly from an FTP or HTTP URL.
{% endhint %}

{% hint style="warning" %}
A *GFF-formatted* file must be paired with a **corresponding FASTA file** of the DNA sequence. These  will be parsed into two data objects: an assembly object with the sequence and a genome object containing the original feature calls and annotations.
{% endhint %}

Further instructions for adding data to your Staging Area can be found [here](/getting-started/narrative/add-data#uploading-data-from-external-sources).

## Importing a GenBank-formatted genome

Using a file on your computer, open the [*Import* tab within the **Data Browser**](/getting-started/narrative/add-data)**.** Then drag & drop the genome file into your Staging Area.&#x20;

Open the *Import As...* pulldown menu to the right of the filename in your Staging Area and select “GenBank Genome.”

Make sure the correct file type is selected and the checkbox is active, then click the “Import Selected” button.

<img src="/files/kzWFK7v4Ozy7grT4VFa2" alt="Import GenBank Genome" width="563">

The Data Browser will close and the “Import GenBank File as Genome from Staging Area” App will be added to your Narrative.

Notice that the name of the Genome file is filled in, as is a suggested name for the Genome and Assembly data objects that will be created by the import, which *can* be changed. Adjust the Genome Type and source of the GenBank file or any advanced parameters if needed.&#x20;

![Import GenBank Genome App](/files/SrZPBrs1f3IwHQbIfKA7)

Click the green "Run" button to start the import. When the import is finished, your **Data Panel** will update to show the new Genome and Assembly objects, and a report will appear in the Import App.

## Importing a GFF-formatted Genome

Open the [*Import* tab in the **Data Browser**](/getting-started/narrative/add-data) and drag and drop *both* the genome *and* corresponding FASTA file into your Staging Area. In your Staging Area, open the *Import As...* pulldown menu to the right of the GFF filename and select “GFF Genome."

<img src="/files/UVeIGpfLgpRRx3ZISggF" alt="" width="563">

Note the name of the corresponding FASTA file.

Make sure the correct file type is selected and the checkbox is active, then click the "Import Selected" button. The data slide-out will close and the “Import GFF/FASTA File as Genome from Staging Area” App will be added to your Narrative. The GFF File Path name will be filled in.

You will need to fill in the name of the FASTA file. Using the dropdown for the “FASTA File Path”, select the FASTA file in the Staging Area. Ensure the file type is a FASTA file type.&#x20;

![GFF/FASTA Genome Import](/files/4nZWhJGScKIBmqnfIAi0)

The name of the Scientific Name may be filled in, as is a suggested name for the Genome data object that will be created by the import. You can edit the name of the output Genome Object Name, Scientific Name, and any advanced options as needed. Click the green "Run" button. When the import is finished, your Data Panel will update to show the new Genome object, and a report will appear in the Import App.

## Uploading a Genome from other sources

You can upload data into your KBase Staging Area using [Globus](/data/globus), or by supplying a URL for a publicly accessible FTP location, Google Drive, Dropbox, or a direct HTTP link to import into the Narrative. Options for adding data to your Staging Area are described [here](/getting-started/narrative/add-data).

<img src="/files/KzHe3U0JdzLl1GIJjBoi" alt="Drag and Drop, Upload with Globus, Upload with URL options to upload and import data into KBase" width="563">

### Bulk Import

Both GenBank and GFF genomes can be imported as one of the supported bulk import types. You can select multiple assemblies simultaneously from the staging area to import them at once. See the bulk import section of [the guide to importing data into the Narrative.](https://docs.kbase.us/getting-started/narrative/add-data)&#x20;

## Uploading a Genome using a URL&#x20;

Open the [*Import* tab in the **Data Browser**](/getting-started/narrative/add-data) and click on the *Upload with URL* button (below the drag & drop area) to open an Upload App Cell.

The Data Browser will close and the “Upload a File to Staging from Web” App will appear in your Narrative. Alternatively, you can open the app directly from the **Apps Panel**. From the app, click on the dropdown for the URL Type and select the URL type.

![Upload File to Staging from Web App](/files/9wT13c3HgUduyOmwYqk4)

{% hint style="info" %}
When uploading a GenBank Genome, you will only need to use one link. When uploading a GFF Genome, you will need to use two links for the GFF file and the FASTA file.&#x20;
{% endhint %}

In the App, click the "+" button for the URLs and paste in the name of the Genome file (GenBank or GFF). Hit the "+" button again and paste in the name of the FASTA file.

Then click the green "Run" button to start the upload. After the App completes the files will appear in your Staging Area, which you can access via the *Import* tab in the **Data Browser**.

The genome file(s) are now in your Staging Area. Now you need to import them to your Narrative to use them in analyses.


# FASTQ/SRA Reads

Formatting and uploading FASTQ and SRA reads files.

In KBase, reads from FASTQ and SRA files can be imported to create reads library data objects. The objects will either be a SingleEndLibrary or a PairedEndLibrary. The tools in KBase can then be used to assemble reads into an “Assembly” data object or to align reads to an “Assembly." After uploading and importing reads data, you may want to refer to the documentation about [Assembly and Annotation](/apps/analysis/assembly-and-annotation). Reads can also be used in [RNA-seq and expression analysis](/apps/analysis/expression).

{% hint style="info" %}
Single-end and paired-end reads can be uploaded in FASTQ or SRA format. For FASTQ files, please ensure that your filename ends with the .fastq, .fnq, or .fq file extension. SRA files should have the .sra file extension. The uploader also accepts compressed files in .zip, .gz, .bz2, .tar.gz, .tar.bz2 formats.
{% endhint %}

Files can be uploaded into your KBase Staging Area from your local computer or directly from a publicly accessible FTP or HTTP URL.

## Importing reads from your computer

### Single-end library in FASTQ or SRA format

Using a file on your computer, open the [*Import* tab within the **Data Browser**](/getting-started/narrative/add-data)**.** Then drag & drop the single-end library into your Staging Area. Open the pulldown menu to the right of the filename in the Staging Area and select “FASTQ Reads NonInterleaved" or "SRA Reads."

<img src="/files/EuSz6AYLPKyP8QLKEczy" alt="FASTQ Single-end Library" width="563">

<img src="/files/4tQgUOHTXhryoytWBjcP" alt="SRA Single End Library" width="563">

Make sure the correct file type is selected and the checkbox is active, then click "Import Selected". The slide-out Data Browser will close and an app called “Import FASTQ/SRA File as Reads from Staging Area” will be added to your Narrative.

Notice that the name of the Reads Library file is already filled in, as is a suggested Reads Object Name that will be created by the import (you can change that if you like). Adjust the Sequencing Technology and any of the advanced options if needed. Note that this was a metagenomic sample, we would *uncheck* the box next to Single Genome. When ready, click the green "Run" button to start the import. When the import is finished, your Data Panel will update to show the new SingleEndLibrary object, and a report will appear in the import app cell.

### Paired-end library in FASTQ format

There are two ways that KBase and [GenBank SRA](https://www.ncbi.nlm.nih.gov/sra/docs/submitformats/) recognize a paired-end library. A paired-end library can be either two files which typically have the same name and are designated as forward and reverse or a single interleaved file. Interleaved files use an 8-line format where forward and reverse reads alternate.

Open the [*Import* tab in the **Data Browser**](/getting-started/narrative/add-data#uploading-data-from-external-sources) and drag either one interleaved file or two paired files into the Staging Area.

Open the pulldown menu to the right of the filename under the *Import As...* column in your Staging Area and select “FASTQ Reads NonInterleaved” for the first file in the pair. Make sure the correct file type is selected and the checkbox is active, then click "Import Selected". The Data Browser slide-out will close and the “Import FASTQ/SRA File as Reads from Staging Area” App will open.

<img src="/files/6R7Tv82AssKyQwcSHSFk" alt="" width="563">

Notice that the name of the FASTA/FASTQ file is filled in. A suggested Reads Object Name is also created by the import and can be changed.

If the files are two paired file, you will need to select a file for “Reverse/Right FASTA/FASTQ File Path” from the dropdown.

![Paired-end or non-interleaved FASTQ Import](/files/USE0Sn8ZnpCUYk1HvezZ)

As with the single-end library example, you can make adjustments to the available parameters. Adjust the Sequencing Technology and any of the advanced options if needed. If the file is an interleaved paired-end library, check the box to the right of Interleaved.

![Interleaved FASTQ Import](/files/1bTRM9V1v5Pjy3qXGGlK)

When ready, click the green "Run" button to start the import. When the import is finished, your Data Panel will update to show the new PairedEndLibrary object, and a report will appear in the Import App.

## Importing a reads library from other sources

In the Staging Area, beneath the box for Drag and Drop, there are other options for adding data to your staging area. You can import reads into KBase using [Globus Online](/data/globus), or by supplying a URL for a publicly accessible FTP location, Google Drive, Dropbox, or a direct HTTP link.

<img src="/files/KzHe3U0JdzLl1GIJjBoi" alt="" width="563">

If your reads are in a publicly accessible URL, you can bypass the Staging Area and directly import reads into your Narrative using one of these three apps (which you can find in the Apps panel or the [App Catalog](https://kbase.us/applist/)):

* [Import SRA File as Reads From Web](https://kbase.us/applist/apps/kb_uploadmethods/import_sra_as_reads_from_web/release)
* [Import Single-End Reads From Web](https://kbase.us/applist/apps/kb_uploadmethods/load_single_end_reads_from_URL/release)
* [Import Paired-End Reads From Web](https://kbase.us/applist/apps/kb_uploadmethods/load_paired_end_reads_from_URL/release)

Note that directly importing these files into KBase from the web behaves the essentially the same as uploading to the Staging Area and then importing to a Narrative, except the transfer is carried out by the Importer App. Use the link available to directly download the file as if you were going to save it to your computer.&#x20;

For SRA reads from NCBI, this generally means navigating to the *Data access* tab of the *Run Browser in the NCBI Sequence Read Archive* (SRA) for the reads to import, as seen here:

![For this example, SRR18272216, navigate to the Data access tab and copy the SRA-download link under Name and paste the full link into the Import from Web App. ](/files/pwTZdy9ocdN8YudkFRYb)

Copy the SRA-download link located under the *Name* heading and paste the link into the URL input within the Importer App.&#x20;

For how to search and download or locate download links for SRA sequences, [see the NCBI Search and Download](https://www.ncbi.nlm.nih.gov/sra/docs/sradownload/) documentation.&#x20;

### Bulk Import

FASTQ and SRA reads can be imported as one of the supported bulk import types. You can select multiple assemblies simultaneously from the staging area to import them at once. See the bulk import section of [the guide to importing data into the Narrative.](https://docs.kbase.us/getting-started/narrative/add-data)


# Flux Balance Analysis (FBA) Model

Data handling to create a model to predict biomass yield given a certain amount of input nutrient.

{% hint style="info" %}
Flux Balance Analysis (FBA) models can be uploaded into KBase as a Systems Biology Markup Language (SBML) file using the .sbml or .xml file extension, an Excel file using the .xls extension, or a tab-separated values (TSV) file using the .tsv extension.&#x20;

When uploading an FBA model in the TSV format, there will be two files corresponding to chemical compounds and reactions.
{% endhint %}

## Formatting Flux Balance Analysis Models

#### SBML

More SBML FBA Models for various organisms can be found here: <http://systemsbiology.ucsd.edu/Downloads>.\
For a description of how to create a COBRA-compliant SBML file that can be imported as an FBA model into KBase, please see [SBML Level Three Specifications](http://sbml.org/Documents/Specifications).

#### Excel

In the Excel format, the tables containing the information about the chemical compounds and reactions in an FBA model are stored in two separate tabs respectively. The most important requirement for the Excel file is naming the tabs “ModelCompounds” and “ModelReactions.”

#### TSV

For the Tab-Separated Values file format, the chemical compounds and reactions tables are saved as two separate files named “FBAModelCompounds.tsv” and “FBAModelReactions.tsv” that will be transformed

### **FBAModelCompounds**

<table data-header-hidden><thead><tr><th width="150">id</th><th width="301">name</th><th width="160">formula</th><th>charge</th><th>aliases</th></tr></thead><tbody><tr><td>id</td><td>name</td><td>formula</td><td>charge</td><td>aliases</td></tr><tr><td> cpd00113</td><td> Isopentenyldiphosphate</td><td> C5H10O7P2</td><td> -2</td><td></td></tr><tr><td> cpd02590</td><td> all-trans-Heptaprenyl diphosphate</td><td> C35H58O7P2</td><td> -2</td><td></td></tr></tbody></table>

* **id**: Compound identifier; see the KBase Biochemistry reference for a list of compounds \[link]
* **name**: Name of chemical compound
* **formula**: Chemical formula using the Hill system
* **charge**: Formal charge of the molecule
* **aliases**: Alternative names for chemical compound

### **FBAModelReactions**

Importing an FBA model into KBase requires specified reaction IDs for the biomass producing equations in the model. An individual organism FBA model needs a single biomass equation, while a community model may have multiple biomass producing equations. In the FBAModelReactions table, denote the biomass producing reaction of an individual organism, or the first in a community model, with an ID such as “bio1” so the reaction is easily identifiable. For subsequent biomass producing reactions in a community model, it is recommended that you use “bio2” for the second biomass producing reaction, “bio3” for the third, and so on.

{% hint style="info" %}
Scroll table below from left to right to see complete format
{% endhint %}

<table data-header-hidden><thead><tr><th width="162">id</th><th width="150">direction</th><th width="150">compartment</th><th width="150">gpr</th><th width="215">name</th><th width="150">enzyme</th><th width="150">pathway</th><th width="150">reference</th><th>equation</th></tr></thead><tbody><tr><td>id</td><td>direction</td><td>compartment</td><td>gpr</td><td>name</td><td>enzyme</td><td>pathway</td><td>reference</td><td>equation</td></tr><tr><td>rxn00001_c0</td><td>-></td><td>c0</td><td>fig|211586.9.peg.3751</td><td>Pyrophosphate phosphohydrolase_c0</td><td></td><td></td><td></td><td>(1) cpd00001[c0] + (1) cpd00012[c0] -> (2) cpd00067[c0] + (2) cpd00009[c0]</td></tr><tr><td>rxn00056_c0</td><td>&#x3C;-></td><td>c0</td><td>fig|211586.9.peg.1038</td><td>Fe(II):oxygen oxidoreductase_c0</td><td></td><td></td><td></td><td>(4) cpd00067[c0] + (4) cpd10515[c0] + (1) cpd00007[c0] &#x3C;-> (4) cpd10516[c0] + (2) cpd00001[c0]</td></tr></tbody></table>

* **id**: Reaction identifier; see the [KBase Biochemistry reference for a list of reactions](https://github.com/ModelSEED/ModelSEEDDatabase/tree/v1.0/Biochemistry)
* **direction**: Directionality of chemical reaction specified as forward (->), reversed (<-), or equilibrium (<->)
* **compartment**: Cellular compartment that the reaction takes place in; see the [Metabolic Modeling FAQ for more information](/workflows/metabolic-models/faq-metabolic-modeling)
* **gpr**: Gene-Protein Reaction in PATRIC identifier format
* **name**: Protein associated with the reaction
* **enzyme**: Optional, can be left empty
* **pathway**: Optional, can be left empty
* **reference**: Optional, can be left empty
* **equation**: Chemical reaction equation specified with compound coefficients in parentheses, compound identifiers, and directional arrows

## Importing FBA Models

In order to successfully upload an FBA model into your KBase workspace, you first need to [add the Genome](/data/upload-download-guide/genome) that corresponds to referenced in the FBA Model you wish to upload. Add the Genome to the Narrative before importing the FBA model file.

Using a file on your computer, open the [*Import* tab within the **Data Browser**](/getting-started/narrative/add-data)**.** Then drag & drop the FBA model file (plus compound file if using TSV format) into your Staging Area. Open the pulldown menu to the right of the filename in the Staging Area and select “FBA Model."&#x20;

<img src="/files/4LoiKHCGP1VofVdYHrQ4" alt="FBA Model selection from the Staging Area" width="563">

Now click the import icon (up arrow) to the right of “FBA Model." The Data Browser slide-out will close and an app called “Import TSV/XLS/SBML File as an FBAModel from Staging Area” will be added to your Narrative.

Notice that the Model file path is filled in, as is a suggested name for the FBA Model data object that will be created by the import (you can edit this).

![Import TSV/XLS/SBML File as an FBAModel from Staging Area App](/files/VdAgxwTnS4FksQXdOlMJ)

At this point, the corresponding Genome needs to be linked to the FBA Model. Add the name of the Genome using the dropdown to select the Genome that you added to the Data Panel. Although specifying the Genome is optional, adding a Genome with gene IDs that match genes IDs in the model is required to import gene-protein-reactions (GPRs). GPRs are necessary to perform gene KO in the Run Flux Balance Analysis App.

Ensure the correct Model file type is selected (SBML, Excel, Tab-separated values). &#x20;

Click the "+" button to the right of Biomass and type in the name of the first biomass. If there are additional biomass values to enter, click the "+" button to open another parameter field for another biomass. *Note: If the biomass reaction name starts with “R\_”, do NOT include the “R\_” when entering the name in this field.* &#x20;

If importing a TSV-formatted model, select the corresponding Compounds file path from the dropdown.

Edit the FBA Model object name for the output if necessary. Click the green "Run" button to start the import.

When the import is finished, your Data Panel will update to show the new FBA Model object, and a report will appear in the Import App.

## Associating an FBA Model with Genome

For KBase Metabolic Modeling Apps, associating a genome with user uploaded models is necessary for ensuring fidelity between the genome annotations and gene-protein-reactions in the metabolic model. This means that the gene IDs between the model and genome need to match. Occasionally, researchers will use locus tags or gene IDs for constructing metabolic models that are different than those found commonly in reference genomes. If the gene IDs do not match, you will need to change either the model or the genome to ensure their gene IDs match.&#x20;

One way to accomplish this is to locate the genome used as the reference for building the model – based on shared gene IDs – and to upload this genome into KBase before uploading the model and associating it with the genome. If you cannot locate the genome with the gene IDs used to build the model, the best option for reconciling the gene IDs is to edit the model file (as SBML or TSV) to change the gene IDs contained therein to match the gene IDs of the genome.&#x20;

This [public Help Board ticket](https://kbase-jira.atlassian.net/browse/PUBLIC-20) has a detailed example of how to accomplish this.


# Media

How to format reaction media files for KBase to use in FBA and modeling analysis.

KBase offers several options for defining specific growth or culture media for metabolic modeling. Within the Narrative Interface, you can use the [Edit Media](https://kbase.us/applist/apps/fba_tools/edit_media/release) tool to build media based on the hundreds of conditions available in KBase or import your own custom media as a TSV or Excel file using the instructions below.

## Formatting Media TSV or Excel files

If you choose to upload a custom media or use media from an external source, please ensure that it is formatted properly for use with KBase.&#x20;

{% hint style="info" %}
Media can be uploaded as a TSV (tab-separated values) or Excel file using four columns:

* **Maxflux** (maximum allowed uptake/excretion of compound)
* **Minflux** (minimum allowed uptake/excretion of compound)
* **Concentration** (concentration of compound in mol/L)
* **Compound identifier** (e.g., ModelSEED ID, KEGG ID, PubChem ID, or compound name such as “glucose”)
  {% endhint %}

A list of all the biochemical *compounds* is available in the [KBase ModelSEED Biochemistry Database ](https://github.com/ModelSEED/ModelSEEDDatabase/tree/v1.0/Biochemistry)for a reference list of reactions and compounds or use the [Biochemistry Search](https://narrative.kbase.us/#biochem-search). From the list you can find the IDs for compounds to include in custom media.

The *concentration* field is not used by any of the the metabolic modeling tools in KBase. It is best to account for these values accurately so that your media files will retain their utility pending updates.

The *minflux* and *maxflux* columns connect media to a flux balance analysis (FBA) metabolic model. The values in these columns are absolute units representing the range of possible fluxes for reactions that transport the media compounds into and out of the cell. A negative flux value corresponds to excretion (transport out of the cell), and a positive flux value corresponds to uptake (transport into the cell).&#x20;

The combination of maximum and minimum flux will dictate whether uptake and/or excretion of a reagent is allowed. Assigning a positive number to *minflux* for a particular compound essentially forces uptake of that compound. Assigning a negative number to *maxflux* will force excretion of the compound.

Minimum and maximum flux are measured in mmol per gram cell dry weight per hour.

Note: When creating a media Excel file, the name of the worksheet that contains your media conditions **must** be named “MediaCompounds” or it will not upload. By default, the first sheet of any Excel file is called “Sheet1.” To rename it, find the worksheet tabs at the bottom left of the Excel file. Double-click on the default name and type in “MediaCompounds.”

Below is an example of a media file in TSV format. In this case, the compound IDs in the first column are from the [KBase ModelSEED Biochemistry Database](https://github.com/ModelSEED/ModelSEEDDatabase/tree/v1.0/Biochemistry) (scroll to the right to see full table):

| compounds | name      | formula | minflux | maxflux | concentration |
| --------- | --------- | ------- | ------- | ------- | ------------- |
| cpd00149  | Co2+      | Co      | -100    | 100     | 0.001         |
| cpd00099  | Cl-       | Cl      | -100    | 100     | 0.001         |
| cpd00063  | Ca2+      | Ca      | -100    | 100     | 0.001         |
| cpd00007  | O2        | O2      | -100    | 10      | 0.001         |
| cpd00027  | D-Glucose | C6H12O6 | -100    | 5       | 0.001         |

## Import media from a TSV or Excel file

For this example, you can click on either link to download a small media file onto your computer:

{% file src="/files/-MJXZqM-GLpO2SnnoaP1" %}
Example Excel file
{% endfile %}

{% file src="/files/-MJXZKO3U5oKJ30MeG5t" %}
Example TSV file
{% endfile %}

Using a file on your computer, open the [*Import* tab within the **Data Browser**](/getting-started/narrative/add-data)**.** Then drag & drop the media file to upload it to your Staging Area. Choose "Media" from the *Import As...* data type dropdown menu in your Staging Area. Click the upload icon next to the field that now says “Media.”

<img src="/files/nIJbrg8RHcRsvQb5qIGb" alt="Selecting Media for import from the Staging Area" width="563">

The “Import Media file (TSV/Excel) from Staging Area” App will be added to your Narrative.

![Import Media file App](/files/juX8BVDUjnfHWdyfnAhc)

You can change the name that will be assigned to the Media Object. Then click the green "Run" button. After the import process has completed, the Media data object will appear in your Data Panel.


# Expression Matrix

The Expression Matrix data type contains gene expression values taken under given sampling conditions.

{% hint style="info" %}
Expression Matrices can be uploaded and imported using the Tab-separated values (TSV) file extension .tsv or the Excel file extension .xls.
{% endhint %}

## **Formatting Expression Matrix files**

If you are importing expression data from an external source or want to populate a file with your own data, please ensure that it is formatted properly for use with KBase.&#x20;

#### TSV

A tab-separated values (TSV) file is a tab delimited text file that has genes across the rows and sample observations across the columns. Make sure the first label in the first column is “feature\_ids” followed by tab-delimited labels for samples.

#### Excel

In Excel, the first column matches *Feature IDs* from the genome. This column could have also been a gene alias. Gene aliases supported by KBase include NCBI, EMBL, UniProt, BioCyc, and ASAP. The column headings are *sample conditions*.&#x20;

The remaining cells in the table contain expression values for the appropriate gene and sample. Be sure to exclude gene features that are missing all expressions or are composed of non-changing expressions across the samples.

Below is an example of a properly formatted expression data file in TSV format. In this case, the gene-ids in the first correspond to gene identifiers for *E. coli* K-12 MG1655 genes and the sample conditions are derived from the [Many Microbe Microarrays Database (M3D)](http://m3d.mssm.edu/about.html).

| feature\_ids | dinI\_U\_N0025\_r1 | dinI\_U\_N0025\_r2 | dinI\_U\_N0025\_r3 |
| ------------ | ------------------ | ------------------ | ------------------ |
| b4634        | 9.05367            | 9.07827            | 9.10114            |
| b3241        | 7.20924            | 7.08695            | 7.07071            |
| b3240        | 7.21535            | 7.14312            | 7.19478            |

To see an example Expression Matrix format, [download](/data/upload-download-guide/downloads) the "Shewanella\_MR-1\_M3D\_ExpData" ExpressionMatrix data set from the *Example* tab in the Data Panel. Unzip the downloaded file and examine the "matrix.tsv" file.&#x20;

### **Plant Expression Data**

For KBase plant genomes, the gene IDs retain the data structure from the external source databases (Ensembl or Phytozome) and do not have aliases as mentioned above. When constructing an expression dataset, append gene IDs with the transcript IDs followed by “.CDS." You can check that you have the correct gene IDs&#x20;

## Upload and Import an Expression Matrix

Expression Matrices can be uploaded into KBase without a Genome, but they will have limited use. In order to successfully upload an Expression Matrix into KBase, you first need to add the [Genome](/data/upload-download-guide/genome) that corresponds to references in the Expression Matrix.&#x20;

Using a file on your computer, open the [*Import* tab within the **Data Browser**](/getting-started/narrative/add-data)**.** Then drag & drop the expression matrix file into the Staging Area.

Now that the expression matrix is in your Staging Area and you can import it from there into your Narrative. Open the pulldown menu to the right of the filename under the *Import As...* column in your Staging Area and select “Expression Matrix."

Click the import icon (up arrow) to the right of “Expression Matrix.” The Data Browser will close and the “Import TSV File as Expression Matrix From Staging Area” App will be added to your Narrative.

The name of the Tab-delimited (TSV) file is already filled in, as is a suggested name for the Expression Matrix data object that will be created by the import, which can be edited or changed.

![Import Expression Matrix App](/files/A93wQbIKmoFMkIJ6qAxA)

At this point, the corresponding genome is optional and hasn’t been linked to the expression matrix. To add the name of the Genome, click on the ‘show advanced’ link to the right of ‘Input Objects’ in the Import App.

![Import Expression Matrix App showing advanced parameters](/files/BvbkcxgaUmyeDld4toRc)

Use the Genome dropdown to select the corresponding genome. Adjust any of the other advanced options if needed, then click the green "Run" button to start the import. When the import is finished, your Data Panel will update to show the new Expression Matrix object, and a report will appear in the Import App.&#x20;

#### Expression Matrix Object

Each gene measured in the expression dataset should have an identifier listed in the first column of the TSV file. To ensure that the gene identifiers listed in your dataset correspond to the aliases in KBase, click the name of the Genome object to open up the genome viewer.&#x20;

* Click on the tab labeled *Browse Features* and locate a gene of interest by searching for the name of the function or protein associated with the gene.&#x20;
* Click the *Feature ID* of the gene of interest to open a new tab with additional information about the gene.&#x20;
* Locate the section titled *Aliases* and crosscheck the gene labels contained within your expression dataset with either the Feature ID or one of these aliases to ensure that these labels will correspond to features in KBase.


# Phenotype Set

The Phenotype Set data type represents experimental data about an organism’s ability to grow in specific media conditions, recorded as growth or no growth.

## **Formatting Phenotype files**

{% hint style="info" %}
A Phenotype Set can be uploaded from a TSV (tab-separated values) file with a .tsv or .tab file extension, or from Excel spreadsheet with a .xls extension. The spreadsheet needs to have exactly these five columns:

* **Gene knockout (geneko)** – List of genes knocked out in the phenotype; use 'none' for wild-type phenotypes. Gene IDs should be in the same format that appears in your metabolic model (e.g., kb|g.220339.CDS.2927)
* **Workspace information (mediaws)** – Workspace where the media for the phenotype data was loaded into KBase. The workspace information can be found by running the command `print(os.environ['KB_WORKSPACE_ID'])`in a code cell. The output should be formatted `username:narrative_#############.`
* **Media** – ID of the media condition loaded in KBase where the phenotype was observed.
* **Additional Compounds (addtlCpd)** – Additional media compound IDs to be added alongside the primary media formulation. See the [KBase ModelSEED Biochemistry Database ](https://github.com/ModelSEED/ModelSEEDDatabase/tree/v1.0/Biochemistry)for a reference list of reactions and compounds or use [Biochemistry Search](https://narrative.kbase.us/#biochem-search).
* **Growth** – Indication of whether or not the organism grew in the specified media with the specified knockouts. 1 meaning growth; 0 meaning no growth.
  {% endhint %}

| geneko   | mediaws    | media       | addtlCpd | growth |
| -------- | ---------- | ----------- | -------- | ------ |
| SO\_0009 | KBaseMedia | C-D-Glucose |          | 0      |
| SO\_0009 | KBaseMedia | C-D-Lactate |          | 1      |
| SO\_0009 | KBaseMedia | C-acetate   |          | 1      |

To download an example phenotype set data table, go to the "Example" tab in the data Data Browser slide-out, add a Phenotype Data Object to the Narrative and download the TSV file.&#x20;

<img src="/files/3C8SSrRoamKf79wHKBLH" alt="Example Data Objects" width="563">

## Importing a Phenotype Set

First [add the Genome](/data/upload-download-guide/genome) that corresponds to referenced in the Phenotype Set. Add the Genome object to the Narrative before importing the Phenotype Set file.

Using a file on your computer, open the [*Import* tab within the **Data Browser**](/getting-started/narrative/add-data)**.** Then drag & drop the phenotype set file into your Staging Area. Open the pulldown menu to the right of the filename in the Staging Area and select “Phenotype Set."&#x20;

<img src="/files/lMfYCuiOxznEv78H9rzw" alt="Importing a Phenotype Set from the Staging Area" width="563">

Click the import icon (up arrow) to the right of “Phenotype Set”. The data slide-out will close and the “Import TSV File as Phenotype Set From Staging Area” App will be added to your Narrative.

Notice that the name of the Phenotype TSV file is already filled in, as is a suggested name for the Phenotype Set data object that will be created by the import, which you can edit.

![Import Phenotype Set App](/files/NzeKYnIAmYYVvYgzrmt7)

Add the name of the Genome, click on the dropdown to select the corresponding genome that you added to your Narrative earlier.

Click the green "Run" button to start the import. When the import is finished, your Data Panel will update to show the new Phenotype Set object, and a report will appear in the Import App.


# Amplicon Matrix

Generating and handling Amplicon Matrices to link to Samples

## Formatting Amplicon Matrix files

Any approach to creating an amplicon matrix should be usable in KBase, but may require additional curation. Amplicon matrices or taxonomic abundance matrices include OTU (operational taxonomic units) and ASV (amplicon sequence variants).&#x20;

We recommend to upload raw counts as analysis ([rarefaction, standardization, and normalization](/apps/analysis/matrix)) will be tracked and provenance maintained on system. There are a number of [tools in KBase to analyze Amplicon Matrices and metadata](/apps/analysis/matrix).&#x20;

{% hint style="info" %}
Amplicon matrices can be uploaded from a TSV (tab-separated values) file with a .tsv file extension.

When uploading, ensure that **rows are taxa** and **columns are samples**. A *column for taxonomy is optional*.

Consensus sequences must also be included as a FASTA file with extension .fa, .fasta. Sequence names **must match&#x20;*****exactly*** with the taxa names in the amplicon matrix.&#x20;
{% endhint %}

Two files are *needed* for uploading amplicon data:

* *FASTA file* with consensus sequences
* *Taxonomic abundance matrix* (OTU or ASV)&#x20;
  * row = taxon
  * column = sample
  * optional column = taxonomy

Additional information to have on hand for parameters include:

* *Sample processing metadata* (SampleSet), i.e. sequencing technology or platform, target gene or region
* *Bioinformatic processing metadata,* i.e. clustering method, quality filtering steps

## Uploading and Importing

Amplicon Matrices can be uploaded into KBase without linking metadata. For full functionality of analysis tools using metadata, first [upload and import the SampleSet](/data/upload-download-guide/sampleset).&#x20;

Using a file on your computer, open the [*Import* tab within the **Data Browser**](/getting-started/narrative/add-data)**.** Then drag & drop the amplicon matrix file **and** the FASTA file with consensus sequences into the Staging Area.

Once the files are in your Staging Area, you can import the data into your Narrative.

Go to the APPS panel and open the Upload category to locate the "Import Amplicon Matrix from TSV/FASTA File in Staging Area" App. Click on the Import Amplicon App to add it to your Narrative.&#x20;

![Import Amplicon Matrix from TSV/FASTA File in Staging Area. Double click to view. ](/files/nB11SCjruwLswWvvQfAx)

In the first section for Input Objects, select the previously imported SamplesSet file from the dropdown menu for linking metadata.&#x20;

Under Parameters, select the Amplicon Matrix or Taxonomic Abundance TSV file using the dropdown and then the corresponding sequences FASTA file.&#x20;

Parameters that *require* *inputs* are the target gene, target subfragment, taxon calling method, i.e. denoising or clustering, and sequence error cutoff.&#x20;

Additional processing metadata, i.e. primer sequences, library kit, sequencing center, the denoise or clustering methods can be input by clicking "show advanced" parameters. While these parameters are not required, they are recommended to enhance workflow documentation and reproducibility.&#x20;

Fill in the Amplicon Matrix Object Name and click the "Run" button.&#x20;

Once your amplicon matrix is imported into the Narrative, you can use the "View Matrix as Table" App to view your new AmpliconMatrix object. &#x20;

{% file src="/files/-MLxg9TfZ9m06tX34ZKp" %}
Example of test amplicon matrix with consensus FASTA
{% endfile %}

The amplicon matrix may be applied to an existing [SampleSet](/data/upload-download-guide/sampleset) and [Attribute Mapping](https://narrative.kbase.us/#catalog/apps/kb_uploadmethods/import_attribute_mapping_from_staging/beta) on-system to facilitate the interoperability of objects in KBase.&#x20;

## Tutorial Webinar

{% embed url="<https://youtu.be/w7y23FC5JUc?t=573>" %}
Importing Samples and Amplicon
{% endembed %}


# Chemical Abundance Matrix

Metabolomics, exometabolite, and chemical abundance data can be integrated with metabolic modeling and flux balance analysis tools in KBase.

## What is "chemical abundance" in KBase?

The name of this data type “chemical abundance” is a broad term that we use to represent a wide array of measurements associated with chemicals. This data type can be used to upload and store diverse types of chemical data in the system such as metabolomics (intracellular and/or extracellular) that is derived based on microbiomes/isolate organism growth experiments etc., computationally predicted compounds, or data collected on the concentration or from elemental analysis. These data could be collected on environmental samples, such as soil, sediment, or water. Currently, the metabolomics data derived from the samples are the most popular data that is uploaded and stored in the system.

Once chemical abundance data matrices are [uploaded](/data/upload-download-guide/chemical-abundance-matrix#uploading-and-importing), they can be analyzed using [KBase Apps for metabolomics](/apps/analysis/metabolic-modeling#metabolomics), such as Escher mapping. Additional statistical analysis of the chemical abundance attribute maps, such as PCA and clustering, can also be performed.

{% hint style="info" %}
A Chemical Abundance Matrix can be uploaded from a TSV (tab-separated values) file with a .tsv or .tab file extension, or from Excel spreadsheet with a .xls extension.

Each Chemical Type can be either a specific compound or element, aggregate (totals), exometabolites (measurements of compounds or elements that are consumed or excreted into the medium).&#x20;
{% endhint %}

{% embed url="<https://www.youtube.com/watch?v=zSAhBEKFt1o>" %}
Chemical Abundance Upload Webinar
{% endembed %}

## Formatting chemical abundance matrices

The [Create Chemical Abundance Matrix Template App](https://kbase.us/applist/apps/GenericsAPI/build_chemical_abundance_template/release) creates an Excel spreadsheet for direct download that can be populated with chemical abundance data. While Chemical abundance data works best and more meaningful when linked with an existing SampleSet in the system, linking a SampleSet is not required (See section Linking SampleSet). While Chemical abundance data works best and more meaningful when linked with an [existing SampleSet](/data/upload-download-guide/sampleset) in the system, linking a SampleSet is not required.

The minimal set of metadata in a chemical abundance matrix includes an ID (unique value) field, a chemical type (aggregate, exometabolite, specific), and one or more of the following:  Compound ID (e.g; ModelSEED, KEGG, ChEBI), mass, formula, inchikey, inchi, smiles, or compound name. Additional metadata such as units are strongly encouraged to provide with proper information that fits your scientific use cases or be kept as ‘unknown’.  (see section "Template Fields Descriptions" for an explanation of each field)  Providing additional metadata may enhance the downstream analysis of use cases for you and other readers.

If a SampleSet exists, it can be applied to the chemical abundance data. Chemical abundance data needs to be formatted to ensure Samples are correctly linked.

Note that linking to Samples is not required, but highly recommended. When linking to using this app, the template will be automatically populated with Sample IDs to ensure the chemical abundance data is properly linked to corresponding Samples in the system.

![Create Chemical Abundance Matrix Template](/files/tuJS6ceA3b7NtlDrQIGo)

This App generates a spreadsheet onto which you can copy your data to ensure it links to the SampleSet when uploaded.&#x20;

When creating the chemical abundance template, there are 10 columns in the default sheet shown in the sheet below. Column headings in italics come pre-filled with validated values to choose form a dropdown.

<table><thead><tr><th>Column Heading</th><th width="212.33333333333331">Description</th></tr></thead><tbody><tr><td>ID (unique value)</td><td>The identifier for the chemical/element/compound or peak value. This can be user-defined but must be unique for each compound.</td></tr><tr><td><em>Chemical Type</em></td><td><ul><li>Specific - Specific compounds in the sample (e.g; TiO2, Tyrosine) </li><li>Aggregate - Measurement of total concentration of individual elements (e.g; Total C, N) </li><li>Exometabolite - Measurements of compounds that are consumed or excreted into the medium by an organism or a microbial community</li></ul></td></tr><tr><td><em>Measurement Type</em></td><td>unknown, FTICR, Orbitrap, Quadrapole</td></tr><tr><td><em>Units</em></td><td>mg/Kg, g/Kg, mg/L, mg/g DW, mM (millimolar), M (molar), % (percentage), Total Weight  %, unknown</td></tr><tr><td><em>Unit Medium</em></td><td>soil, solvent, water</td></tr><tr><td>Chemical Ontology Class</td><td>Used for grouping chemicals by their functional groups (e.g; aromatics) or could use according to ontologies defined in public databases (e.g; ChEBI).</td></tr><tr><td><em>Chromatography Type</em></td><td>unknown, HPLC, MS/MS, LCMS, GS</td></tr><tr><td>Chemical Class</td><td>Used for defining major classes like lanthanides</td></tr><tr><td>Protocol</td><td>Any term or sentence defining the protocol used in your lab</td></tr><tr><td>Identifier</td><td>An optional identifier/abbreviation that you use for the compound and or element (e.g; Menaquinone => MK-8, Fatty acid => (Iso-C16:1) )</td></tr></tbody></table>

Select chemical data to include, such as aggregate M/Z, compound name, predicted formula, and more, depending on what data you have for upload.&#x20;

| Column Heading                  | Description                                                                                                                                                                                                                                                                                                                                                  |
| ------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Aggregate M/Z                   | M/Z represents the ratio of mass (M) divided by net charge (Z) of the chemical.                                                                                                                                                                                                                                                                              |
| Compound Name                   | Optional name for the compound                                                                                                                                                                                                                                                                                                                               |
| Predicted Formula               | Formula of the compound (if available). The formula could be a predicted formula based on techniques such as Fourier-transform ion cyclotron resonance (FTICR) mass spectrometry, computationally predicted formula based on molecular weight, derived from instruments based on known protocols, or based on the chemical structure determination from NMR. |
| Predicted Structure (smiles)    | Predicted structure using Simplified Molecular Input Line Entry System (SMILES) notation                                                                                                                                                                                                                                                                     |
| Predicted Structure (inchi-key) | Predicted structure using International Chemical Identifier (InChi) key notation.                                                                                                                                                                                                                                                                            |
| Theoretical Mass                | Theoretical molecular mass of the compound                                                                                                                                                                                                                                                                                                                   |
| Retention Time                  | The retention time of the chemical; not validated, so users can include units or not depending on needs                                                                                                                                                                                                                                                      |
| Polarity                        | Polarity of the chemical.                                                                                                                                                                                                                                                                                                                                    |

Finally, select at least one form of standard Chemical IDs to include in the template. These can be selected from the [KEGG](https://www.genome.jp/kegg/compound/), [ChEBI](https://www.ebi.ac.uk/chebi/), and [ModelSEED](https://modelseed.org/biochem/compounds) databases. You can use their respective websites to get the identifier for your compound. If you are creating a chemical abundance sheet from scratch, you don't have to include one of these chemical IDs, but it is recommended that you do so in order to compare similar chemical once uploaded.

{% hint style="info" %}
If you plan to use your data with metabolic modeling analysis pipelines (see Use cases section) we highly encourage to have at least one type of identifier to be listed (if available), as the compound identifiers will be used to map compounds in metabolic models. (Alternatively, you can provide inchikeys which we able to map to compounds in metabolic models).
{% endhint %}

## Uploading and Importing

&#x20;For full functionality of analysis tools using metadata, first [upload and import the SampleSet](/data/upload-download-guide/sampleset).&#x20;

Once you have added and [formatted](/data/samples/ontology) your data to the Chemical Abundance Matrix template, you can upload it using the [Import Chemical Abundance Matrix from CSV/Excel/TSV File in Staging Area App](https://narrative.kbase.us/#catalog/apps/GenericsAPI/import_chemical_abundance/release).&#x20;

Using a file on your computer, open the [*Import* tab within the **Data Browser**](/getting-started/narrative/add-data)**.** Then drag & drop the chemical abundance matrix file into the Staging Area.

Once the matrix is in your Staging Area, you can import the data into your Narrative.

![Import Chemical Abundance Matrix](/files/ytL2jhNTYhe86EJE5PsN)

In the first section for Input Objects, select the previously imported SamplesSet file from the dropdown menu for linking metadata.&#x20;

Under Parameters, select the Chemical Abundance Matrix using the FilePath dropdown.&#x20;

Fill in the Matrix Object Name and click the "Run" button to add the metabolomics data to the Narrative.

![](/files/rMRR03g8uMhIKpVyxT0n)

## Using the Uploaded Data

Once you've uploaded your chemical abundance data, you can explore apps using that type of data or use one of our use case demonstration Narratives. See [the chemical abundance section of in Using Apps for more info.](https://docs.kbase.us/apps/analysis/chemical-abundance)


# SampleSet

A guide to importing and formatting sampling data and metadata to comply with standard templates.

{% hint style="info" %}
SampleSets are still a work in progress, and much of the documentation here is not currently applicable in production. Watch your inbox for notifications from KBase Outreach when these workflows go live.&#x20;

If you are interested in beta testing these Apps, please contact us at <engage@kbase.us>.
{% endhint %}

A Sample in KBase represents physical material from an experiment with attributes or metadata (environment, location, temperature, etc). It has a unique ID internal to KBase, aliases, related Samples, and structured metadata. Samples occur within the Narrative as a SampleSet, and multiple SampleSets may share the same Sample.&#x20;

## Formatting SampleSets

### IGSN SESAR

The [SESAR](https://www.geosamples.org/) (System for Earth SAmple Registration) format is one format that can be used in KBase for uploading Samples. Their samples use a [controlled vocabulary for fields of data and metadata](https://www.geosamples.org/vocabularies) that can be used in KBase.

Download a template of the SESAR format below and fill in your own data, or generate a new dataset file [based on their documentation](https://www.geosamples.org/resources/help). The .xlsx file has pre-filled options for the columns with [controlled vocabulary](/data/samples/ontology), whereas the CSV file is a much smaller file that only contains a heading with the possible fields.

{% file src="/files/-MaU7PJn9LVy3gilLlCz" %}
Downloadable SESAR Template File
{% endfile %}

{% file src="/files/-Mb6xSWXt4aeih41Wmzk" %}
Downloadable SESAR CSV Template
{% endfile %}

Samples can be directly imported from IGSN SESAR using the "Import Samples from IGSN" App, which takes SESAR number(s) separated by commas as input and imports the sample information directly from SESAR.&#x20;

### NCBI BioSamples

NCBI BioSamples can be directly imported using the "Import Samples from NCBI" App, which takes the NCBI BioSample accession number(s), again separated by commas when importing multiple samples into the same SampleSet.

### ENIGMA

The Ecosystems & Networks Integrated with Genes & Molecular Assemblies (ENIGMA) Science Focus Area (SFA) partners with KBase for new functionality development. The ENIGMA metadata template is available as a template for the "Import Samples" App.&#x20;

{% file src="/files/-Me6CVeoAGmgK6WjgNj1" %}
Downloadable ENIGMA XLSX Template
{% endfile %}

{% file src="/files/-Me6Caair52h-laP65SS" %}
Downloadable ENIGMA CSV Template
{% endfile %}

## Uploading and Importing

When uploading and importing Samples, new Samples can either be added to an existing SampleSet or used to create a new SampleSet. To add Samples to an existing SampleSet, choose the appropriate SampleSet from the drop-down in the importer. To create a new set, leave it blank and a new SampleSet will be created from the newly imported Samples.

#### *from IGSN or NCBI*

Samples can be directly imported from IGSN or NCBI using the "Import Samples from IGSN" or "Import from NCBI" Apps, respectively, searchable in the Upload section of the Apps panel. Samples imported using this method can be added to an existing SampleSet through selecting it from the SampleSet drop-down or can be used to create a new SampleSet. Additional parameters are the same as those for importing Samples manually (from local computer).

#### *from local*

Samples data is uploaded to the Staging Area and then imported into a Narrative. When importing new Samples, the default action is to create a new SampleSet. Individual samples must belong to a SampleSet, even if that SampleSet contains only one Sample.&#x20;

Samples must be in [SESAR or ENIGMA formats](#formatting-samplesets) to not limit analysis as these standards are required to validate most of the data. Non-validated data can be uploaded as user terms, but this will not allow comparisons across Samples. We recommend to change the column headers to correspond with the relevant SESAR or ENIGMA headers, and ensure the data matches.

### [Data Validation](/data/samples/ontology)

SampleSets that are uploaded to KBase are validated to ensure the data format is correct. If data does not match format requirements, the upload may result in an [error](#error-report-display).&#x20;

{% hint style="warning" %}
Critical errors, such as unrecognized units for a measurement, will prevent upload entirely. Minor errors, such as an unrecognized controlled vocabulary term, can be ignored by selecting the "Ignore Warnings" checkbox in the importer.
{% endhint %}

### Viewing

#### in the Narrative

Once imported, the SampleSet viewer widget allows you to see the Samples data that has been imported. This includes all metadata, such as collection information, and can be selectively viewed using the search bar.&#x20;

#### the Landing Page

More details for the SampleSets can be found on the landing page. Click on the SampleSet name in the viewer widget or the binoculars button in the data pane to view the landing page.

{% embed url="<https://youtu.be/PFWNWji8p0U?t=461>" %}
Uploading and Importing SampleSets
{% endembed %}

## Permissions

Permissions are managed separately from the Narrative workspace with some exceptions. The permissions for a Sample can be viewed in the **Sample landing page** under the **Access tab.** When a SampleSet is shared with a user, that user automatically has read access to  all Samples included in the SampleSet. Read access will automatically be given with the first access to the SampleSet by visiting its landing page or viewing it in the Narrative.

![Viewing Samples Access Permissions](/files/-McoKLOFqZ8DAOqndN-a)

To grant a user increased permissions (e.g. to modify Samples), you should use the **Update SampleSet Access Controls App** in the Narrative.&#x20;

![Update SampleSet Access Controls App to change user permissions for specific SampleSets. ](/files/GIXYg14L78oCveJsRzlc)

{% hint style="warning" %}
Permissions will be granted for all Samples in the SampleSet. There is not currently a mechanism to change the permissions on a single app and there is not a method to easily revoke permissions.
{% endhint %}

## Updating Attributes

You will need to export a CSV file of the Samples from a SampleSet, modify the columns and/or their values, then re-import the Samples CSV using the Samples Importer App to edit Sample attributes (for example adding, modifying or removing columns or updating values).&#x20;

This will not modify existing data or analysis that has already been run. To update the analysis results, you will need to re-run the analysis apps using the new SampleSet with your recently updated Samples.

### Subsets and Supersets

Creating subsets and supersets in the Narrative is under development. The most efficient way manage subsets of Samples is to upload separate SampleSets for each subset, and a complete SampleSet to represent the parent set. In the same way, a superset can be created by downloading SampleSets, combining them, and then uploading the combined superset.

## Errors

The Samples Importer Apps are designed to encourage the use of validated data, but we are aware that not all information is accounted for in our validator, and sometimes wrangling the data to fit validation formats is not worth the effort. While we do encourage validation as much as possible, our importers use a two-tiered system of error reporting.

There are two, color-coded error types: red and yellow. &#x20;

1. Red errors indicate the validation failed to recognize a value as a validated field, such as the type "playdoh" for material. This prevents the importer from creating a SampleSet for any cell with values that do not comply with controlled vocabulary.&#x20;
2. Yellow errors indicate a column (in this case, "color") could not be validated as a controlled field. These result in warnings, but the SampleSet can still be created.

![Two kinds of errors - red errors are critical, while yellow are warnings](/files/-MePywAXeYunU5HVOT0B)

### Handling Errors

Sample imports can return a "successful" job that does not create an object in the Data panel. In this case, you can interpret a successful job run to indicate that the sample file was read, but the contents may or may not be valid.

* A cell in a validated column that does not comply with the [controlled vocabulary](/data/samples/ontology) will fail to upload. Consult the [format guide](#formatting-samples-and-samplesets) or source of the template to determine if the value is accepted and if the vocabulary contains a valid term for your value. If your value is not part of the controlled vocabulary and there is not a suitable alternative, you can change the column header to a user defined field.

{% hint style="info" %}
&#x20;If you would like your term added to the controlled vocabulary, [please contact us](/troubleshooting/support).&#x20;
{% endhint %}

* If there is a warning due to an unvalidated field, you can choose to accept the warning or decline it and fix the spreadsheet. If you accept an unvalidated field, it will be stored as a user-defined field. This allows you to include the information, but does not allow comparisons across Samples. For additional descriptive information, accepting the warning may be beneficial. But if the information is critical to Samples which you plan to use in analysis, we recommend manually fixing the values to comply with controlled vocabulary terms.&#x20;
* No mechanism exists to accept warnings *after* the import has run. You can run the Importer App with the "Ignore Warnings" option unchecked. This allows you to see the errors and warnings and run does not create a data object. Then proceed to edit your spreadsheet and re-run the App with the warnings fixed, or run the App with the option checked and delete the resulting object if you find warnings you want to fix.


# Compressed/Zipped Files

How to handle compressed files when uploading and importing data.

### Compressed Format Types

Data importers accept compressed files in these formats: .zip, .gz, .bz, .bzip2, .bz2, .tar, .tgz.&#x20;

### Uploading and Importing Compressed Files

To add compressed data files from your computer to the KBase Narrative, follow the same three-step process for all data files.&#x20;

1. Drag & drop the data file(s) from your computer to the new Import tab to upload them to your Staging Area.
2. Choose a format for importing data from your Staging Area into your Narrative.
3. Run the Import App that is created.

When files have been uploaded to the Staging Area, compressed files have an outward-facing double arrow icon to the left of the file name.&#x20;

Click the uncompress icon to the left of the file name to uncompress the compressed/zip uploaded files before clicking the import icon. To ensure this works, the file extension must match the type of packing/compression.&#x20;

Note that importers will accept compressed files, but not archives, as imports. For instance, you can import a single reads file as reads.fastq.gz or decompress it in the staging area and import reads.fastq; however, if you have a series of files in an archive (e.g. reads1.fastq, reads2.fastq,  and reads3.fastq inside all-reads.zip), the archive will first have to be unpacked and then the files can be imported. This applies to both single and bulk imports.


# Bulk Import Specification

Import multiple files at once through bulk upload using an import specification file with the extension .csv, .tsv, or .xls following the instructions below.

{% hint style="warning" %}
Import Specification applies to [select data types](/data/upload-download-guide/uploads#what-data-types-does-this-apply-to).&#x20;
{% endhint %}

{% hint style="info" %}
A TSV (tab-separated values) file with a .tsv or .tab file extension, a CSV (comma-separated values) file with a .csv extension, or an Excel spreadsheet with a .xls extension can be used to upload multiple files at once. &#x20;

This approach treats the Import Specification file as a directory or manifest containing the upload information for multiple files to fill out an import cell. The TSV, CSV, or Excel file must be formatted exactly to import correctly.&#x20;
{% endhint %}

#### Import Specification files (TSV, CSV, Excel) require *several* columns for uploading sequencing data when selecting files. These will vary for each data type, but make sure you have the following information:&#x20;

* **File Path(s)** – These will be relative to your home directory of the Staging Area, so you only need to include the file name if the files are at the top level. If they are within a folder called “folder1” then the file input fields should be "folder1/file1", "folder1/file2", etc.
* **Object name** – Designate an object name for the imported file, which will be the object name to use as the input for analysis Apps.
* **Parameters** – Any required input parameters.&#x20;

{% hint style="info" %}
Use the bulk Importer App with a subset of files to generate an Import Specification template. Download and modify the template locally to fill out file paths and parameters. Then upload the Import Specification file to import all files at once.&#x20;
{% endhint %}

## Creating an import specification template from the Narrative

To create an import specification template CSV, TSV or Excel file for many files from the Narrative, begin with the Bulk Import directions using a subset of the files or file types to import.&#x20;

Drag and drop all files or a folder containing the files to upload to Staging.&#x20;

Once the files appear in the Staging Area, select at least one that represents each file type you wish to import, select the **Import As** file type and make sure the check box is active, then click "Import Selected."&#x20;

This will create a bulk Importer App in the Narrative. Fill out the parameters for the import. Then click the "Create Import Template" button. &#x20;

![Import from Staging Area Bulk Import App](/files/Od4q2OgWBnpCgfP8QM2Y)

This will create a pop-up file to prompt the template creation. Select the import type(s), if multiple use key controls (such as command-hold and select with Mac or control-shift and select on Microsoft). Choose the output type as either Comma-separated (CSV), Tab-separated (TSV), or Excel (XLS). Then select or rename the output destination within the Staging Area. Click "Generate template!" to create the templates selected. &#x20;

{% hint style="warning" %}
If you have an Import Specification Template within the destination that you designate in the pop-up, it will be overwritten with the newly created template.&#x20;
{% endhint %}

<img src="/files/jZsnlWH2XXSUF2MA66qS" alt="" width="563">

{% hint style="info" %}
To import multiple data types at once, either create a CSV/TSV for each type, or a tab for each type in a single Excel file.&#x20;
{% endhint %}

A pop-up will confirm the Import Specification template has been generated and its location in the Staging Area.&#x20;

<img src="/files/MTM2O7MYgOe774ZpYTYj" alt="" width="563">

Navigate back to the Staging Area  and click "Refresh" if you do not see the folder or template in the destination designated in the previous step. Download the Import Specification template using the download button on the left of the trash icon.

<img src="/files/XiF0NakpBCGhb7Wog2rl" alt="" width="563">

Open the Import Specification template. In the template files 3 rows are already filled; these are required for the staging service to parse the file. In the Excel file, the top two rows are hidden to simplify the view. The third row displays the column headers for inputs, outputs, and parameters included in the Importer App. The following rows are for each of the files to import, the files selected before will be included here.&#x20;

Continue filling out the spreadsheet as if you would the Importer App with the remaining files to import. Each row corresponds to one data object to import. Some data objects will include more than one file path. For instance, a GFF Genome requires both GFF and FASTA files. All text fields should match exactly to the display text in the Importer. Use the pre-filled parameters to fill out the rest of the rows with designated file paths.&#x20;

{% hint style="info" %}
To manually fill in the Scientific name into a template, navigate to an Import cell with a GenBank Genome Data type. Under Parameters, go to the 'Scientific name', search for the genus/species name you want and select it. Click the copy button on the right. You can now paste the name into the template.&#x20;
{% endhint %}

![Search for the Scientific name where the orange arrow is located, then copy using the icon to the right with the orange circle.](/files/2ChLdbcTlyuf1P8cw6Ek)

## Importing a bulk Import Specification file

Open the [*Import* tab within the **Data Browser**](/getting-started/narrative/add-data)**.** Then drag & drop the Import Specification file and the files to be imported if they are not already in the Staging Area.&#x20;

<img src="/files/-MK5wLNJPCRC_f9h0FpF" alt="Drag and Drop a CSV file" width="563">

Once the file appears in the Staging Area, select "***Import Specification***" from the dropdown in the *Import As...* column. Make sure **only the Import Specification file is selected** with a filled checkbox (do not select the independent files). Click ***Import Selected*** at the bottom right corner. &#x20;

<img src="/files/9rj0JjA7Gj3dcKwd3iWL" alt="" width="563">

Once the Import Specification file has been read, it will create a bulk Importer App in your Narrative and you can proceed as usual following[ the bulk import](/data/upload-download-guide/uploads#bulk-import-guide) directions. Verify the file names and import settings within the App. Click "Run" on the upper lefthand corner of the App to import the files from the Import Specification into the Narrative.&#x20;

![](/files/jsAkQXW6l05JAiZDdlpV)

Once the Importer App has run, the App will signal that the files have been imported or any errors that occurred across individual jobs.&#x20;

### Limitations

Currently, 500 rows or data objects per import is supported with good performance. The current limit is 10000 rows.&#x20;

If you upload an invalid file, you will receive an error that lists the problem encountered. For errors that can be fixed through the interface, the Importer cell will be created but you must correct the error before running the import.

At this stage of development only a single set of parameters can be used per import job. While we plan to add multi-parameter support in the future and have made the Import Specification templates forward-compatible, the import will apply parameters from the first row to all files.&#x20;

{% hint style="info" %}
*Troubleshooting tips:*

* Do not change the first 3 rows of the TSV or CSV file or the first row or hidden rows of Excel files. This will cause errors or incorrect data.&#x20;
* Be careful of stray cells in Excel - ensure that all the cells outside your data block are empty. Consider deleting all the rows and columns outside your data block.&#x20;
* For CSV and TSV files, be careful to always include the correct number of separator characters for every row, even if some values are blank. The importer is picky about this to help the user avoid off-by-one (or more) errors in their data.&#x20;
* Only one file or Excel tab is allowed per data type. Do not include rows for the same data type in more than one Excel tab, more than one file, a tab and an Excel file, etc.&#x20;
* In CSV and TSV files, white space around non-numerical data is ignored, and whitespace prior to numerical data is ignored.&#x20;
* In CSV and TSV files, you can use quotes (") to surround text that contains the separator character (respectively a comma or a tab).
* If individual files in an import specification are selected along with the import specification, a bug will occur where extra rows that cannot be filled in are added to the bulk import cell for the individual files. To workaround this issue, either do not select files that are within the selected import specification or delete the uncompletable row(s) from the import cell.
* If the scientific name lookup ignores spaces, the workaround is to create (and optionally delete) a new bulk import cell for the scientific name input.&#x20;
  {% endhint %}


# Downloading Data

After analyzing data in KBase, you may wish to download results to your local computer. The Narrative Interface offers the ability to download many of the data types stored in KBase in common form.

## **From the Narrative**

{% embed url="<https://youtu.be/pEKYZbVkyCc?t=329>" %}

* Locate the data you wish to download in the **Data Panel** under the *Analyze* tab
* Hover your mouse over the data object until a “…” icon appears to the right of its name
* Click the “…” icon to reveal advanced options
* Click the "Export / Download Data" icon
* Select a format for exporting and downloading the data
* The data will download to the default downloads folder on your computer

<figure><img src="/files/ALnhgJQkqnq0yA9sO0dF" alt="" width="337"><figcaption></figcaption></figure>

When downloading data from KBase, the data will be compressed into a Zip file (.zip) containing files (or a directory containing files) in the format you selected and a metadata file (in JSON format). If you choose to re-upload data you have downloaded from the system, take note that you cannot directly import the Zip file – you must first extract the file and then re-upload using the specified uploader for the data type.

### **From App Cells**

A few KBase Apps have file links as part of the output at the bottom of the App Cell.

Right click on the file name to download and choose “Save Link As" to download and save to your local computer.&#x20;

## **Large Data Objects**

Large data objects (over 10 gigabases) include many Reads, Assemblies, Genomes, and Alignment (BAM) objects. They encounter many of the same obstacles as uploading large data sets. Just like uploading, this is a 2-stage process; copy the file to the Staging Area and download from Staging to your local computer with [Globus](/data/globus).

*Copy the file to the Staging Area (version 1)*

* Locate the data you wish to download in the **Data Panel** under the *Analyze* tab
* Hover your mouse over the data object until a “…” icon appears to the right of its name
* Click the “…” icon to reveal advanced options
* Click the seventh icon from the left "Export/Download Data"
* Select "Staging"

This will open “Export Data Object to Staging Area” App.&#x20;

The name of the Input Object is filled in. A default Destination Directory is filled in, but can be edited and changed. A directory will be created in the Staging Area, and all created files will be copied to the directory. The "show advanced" Parameters may have more options relevant to your data object.

Once in the Staging Area, large files (>10GB) should be transferred to your local machine using [Globus](/data/globus). Instead of “To” the “KBase Bulk Share” endpoint, this will be a transfer “From” the “KBase Bulk Share” endpoint.

*Copy the file to the Staging Area (version 2)*

* In the **App Panel**, search for the app “Export Data Object to Staging Area,” put in the name of your Input Object for the data object to download, and follow the prior instructions.


# Searching, Adding, and Uploading Data

Video tutorial on searching for, adding to Narratives, and how to upload data in KBase.

{% embed url="<https://www.youtube.com/watch?v=g7-iCVaAMrg>" %}

You can add data to a Narrative through a variety of methods. The **Data Browser** allows you to search data in KBase or import data from your computer or a database.

* *My Data* shows your data objects available within Narratives.&#x20;
* *Shared With Me* includes data associated to Narratives shared with you.&#x20;
* *Public* displays Public datasets.
* *Example* shows datasets that have been pre-loaded by the KBase team.&#x20;
* *Import* allows you to upload your own datasets or external datasets for analysis.

Add data by navigating to the Data Browser tab and hover over the data object to reveal the *<* *Add* button. Click on the "*< Add*" button to import the data object into the Narrative. Now you can view the data object and run analysis.&#x20;

![Add Data already in KBase. ](/files/M3gqfDYiHVFaCle4SvIA)


# Filtering, Managing, and Viewing Data

Video tutorial on filtering, managing, and viewing data in KBase.

{% embed url="<https://www.youtube.com/watch?v=whWUicUWdrY>" %}


# Linking Metadata

How to link environmental sampling metadata with experimental data.

{% hint style="info" %}
SampleSets are still a work in progress. Watch your inbox for notifications from KBase Outreach when these workflows go live.&#x20;

If you are interested in beta testing these Apps, please contact us at <engage@kbase.us>.
{% endhint %}

A Sample in KBase represents measured material from an experiment. It has a unique ID internal to KBase, aliases, related Samples, and structured metadata. Samples occur within a Narrative as a SampleSet, but the same Sample may be shared by multiple SampleSets.&#x20;

One key feature is linking the same data across many users and analyses. To do this, Sample aliases allow the same data uploaded by different users to be linked and connect different analyses of the same data. These aliases can also be used to link Samples in KBase with other methods of sample registration, such as ESS-DIVE or SESAR.&#x20;

Because the metadata about a sample is critically important for, KBase provides the option to [upload metadata spreadsheet](/data/upload-download-guide/sampleset)s to link Samples for more detailed analysis.

{% embed url="<https://youtu.be/PFWNWji8p0U?t=1536>" %}
Managing Linked Data
{% endembed %}

### During upload and import

Once a [SampleSet](/data/upload-download-guide/sampleset) is in system, it can be used when uploading [amplicon](/data/upload-download-guide/amplicon-matrix) or [chemical abundance](/data/upload-download-guide/chemical-abundance-matrix) matrices into KBase. This allows the amplicon and chemical abundance analysis to remain linked to the sample metadata and enrich analysis.

When uploading and importing either an amplicon matrix or chemical abundance matrix, select the corresponding SampleSet from the dropdown menu for the input before running the Import App.

### Additional Data Links

Additional existing data can be linked to samples. The app "Link Workspace Objects to Samples" provides a graphical interface for creating the links. In this app, you can select the SampleSet to create links from, then choose the Sample from the set and object from the Narrative. You can add multiple Samples and objects by clicking the + button.&#x20;

If you have a large number of links to create, it would be advisable to use the "Batch Link Workspace Objects to Samples" app.&#x20;


# Ontologies and Validated Terms

Make the most out of your metadata by using controlled vocabulary

**TL;DR: Use vocabulary that can be identified in KBase so your metadata is functional.**&#x20;

When uploading Samples to KBase, there are many metadata terms that are built in and can be validated. Doing so allows your data to be more easily compared with other data and allows us to integrate data when different sources may use different conventions.&#x20;

{% hint style="info" %}
Many standard terms, such as NCBI BioSamples and IGSN SESAR, aim to describe sample metadata. These standards often overlap and may differ significantly depending on the focus of the organization or institution that manages the samples. At the same time, many labs and smaller institutions have their own templates and organization of their data.&#x20;

Thus, there must be a way to integrate similar metadata from different sources and account for the differences between the sources.&#x20;
{% endhint %}

When Samples are uploaded into KBase via a spreadsheet, they are transformed into an internal representation. Columns within the spreadsheet are mapped and transformed based on the template used. Columns that correspond to recognized terms are validated to ensure they properly formatted, such as checking that the value is the proper type (e.g. string versus a number), in the correct range, match an enumerated list, or appear in an ontology.&#x20;

Controlled terms are useful both because they undergo this validation and they provide a more precise meaning for the value and comparisons accounting for units.&#x20;

When the uploader encounters terms it doesn’t recognize, those terms and values will be stored in a user section of the samples. These values still serve a purpose and can be used in analysis within that data set and they can provide contextual information that the original uploader understands.  However, unrecognized terms can not reliably be compared across SampleSets and other samples in the system. For example, two projects may use the same term to represent different concepts (e.g. depth below sea-level or depth below surface).&#x20;

Using templates and controlled terms clarifies the exact meaning of a term and enforces additional validation. For a full list of terms KBase recognizes, see the [validated metadata table](#validated-metadata) below. If there are terms you would like added, please contact us at <engage@kbase.us> to suggest additions.

## Ontology Landing Pages

The Ontology API supports multiple ontology systems. Currently supported systems are Gene Ontology (GO) and Environmental Ontology (ENVO). You can see more information on both ontologies from their own home pages, or view the KBase landing page for a given ontology page with the links and URL format below (respectively).

* [Gene Ontology (GO)](<http://geneontology.org >)
  * KBase: <https://narrative.kbase.us/#ontology/term/go\\_ontology/GO:#######>
* [The Environmental Ontology (ENVO)](https://sites.google.com/site/environmentontology/home)
  * KBase: <https://narrative.kbase.us/#ontology/term/envo\\_ontology/ENVO:########>

## Validated Metadata Terms

The file TSV file below contains a list of the metadata terms currently supported for validation along with a description that provides general formatting direction. You can view the[ most up-to-date version of this list](https://github.com/kbase/sample_service_validator_config/blob/master/metadata_validation.tsv) in the KBase Samples GitHub.&#x20;

{% file src="/files/-Mb6ysgJ7QIMCxShvnRJ" %}
Validated Metadata TSV
{% endfile %}

{% hint style="info" %}
If you have Sample Attribute columns that are represented as validated terms in SESAR, ENIGMA, or KBase formats, you may request the terms to be added to KBase's vocabulary.&#x20;

KBase will review newly proposed attributes to assess: overlap with current supported attributes; existing or relevant ontological frameworks; and presence of and interoperability of units for measured data.&#x20;
{% endhint %}

### Unit Conversions

You can see a full list of the units and their defined relations to each other in [GitHub](https://github.com/hgrecco/pint/blob/master/pint/default_en.txt).

### Terms

To determine if your term is valid, the easiest way is to search the[ ENVO Database](https://www.ebi.ac.uk/ols/ontologies/envo).&#x20;

## Commonly used units:

**Length**

* Base: meter, m
* Micron/micrometer: µm or um, 1\*10^-6 m
* Millimeter: mm, 1\*10^-3 m
* Centimeter: cm, 1\*10^-2 m
* Kilometer: km, 1\*10^3 m

**Time**

* Base: second, s
* Minute: min, 60 s
* Hour: hr, 3600 s or 60 min
* Day: d, 24 hr
* Year: yr or a, 365.25 d

**Mass**

* Base: gram, g
* Microgram: µg, 1\*10^-6 g
* Milligram: mg, 1\*10^-3 g
* Kilogram: kg, 1\*10^3 g


# Public Data in KBase

KBase provides users with a unified resource for analyzing a range of public data together with the data generated from their own experiments. The data in KBase, which includes public data from various public sources as well as internally-generated analysis results, consists of an extensive set of prokaryotic annotated genomes; a selected set of eukaryotic genomes; and thousands of biochemical compounds and reactions.

Please review [KBase's Data Policy](https://www.kbase.us/about/terms-and-conditions-v2/#data_policy) and the sources of our public reference data.

To search public data in KBase, use the [*Public* tab in the **Data Browser** ](/getting-started/narrative/explore-data) within the Narrative Interface.

<figure><img src="/files/2sMxJnuzMKowpdIceB47" alt="" width="563"><figcaption></figcaption></figure>


# Transfer Data with Globus

A how-to on large file transfer into KBase.

Transfer large files (greater than 2GB) to KBase using [Globus](https://www.globus.org/). Drag & drop from your local computer works for many files, but there is a size limit dependent on your computer and browser.&#x20;

Globus is a data management and file transfer system that can facilitate bulk transfer of data (large data files or a large number of files) between two endpoints. The endpoints that apply here are KBase, JGI, and your local computer. The KBase endpoint is called “KBase Bulk Share,” and JGI has their own way to link to Globus. To do any transfer using Globus, you will need a [Globus account](https://www.globusid.org/create).&#x20;

Your local computer can become a Globus endpoint with a [Globus Connect Personal Endpoint](https://docs.globus.org/how-to/). This is necessary for transfers to or from your local computer.

Once you have a Globus account, link to your KBase account to facilitate file transfers. See [Linking Accounts](/manage-account/link-accounts) for further guidance.&#x20;

### Video Tutorial on Using Globus&#x20;

{% embed url="<https://youtu.be/jQJa49OEx7g>" %}
A quick 4-min video on using Globus with KBase
{% endembed %}

### Starting a Globus Data Transfer

![](/files/-M9PjO7_nSKlpuIKvoww)

To initiate a transfer using Globus, click 'Or upload to this staging area by using Globus Online' in the *Import* tab of the **Data Browser** below the Staging Are&#x61;**.**

<img src="/files/-M7tOi4J47bPJGTIHnyq" alt="" width="563">

This brings you to a page that looks like this:

![](/files/-M7tOlp9UwiROcjL_X8C)

The panel on the right should be set to “KBase Bulk Share,” which directs to the Staging Area. In the example image above, the share point is at the root directory, and the *user* needs to click on their individual directory. If needed, Globus can be used to create new directories in your Staging Area for your data transfer and to delete staging files when you are done importing into KBase.

![](/files/-M7tRDV_jU1Y2RYJ7qTk)

Set the left panel to the location of the data you want to transfer (i.e. you could choose the JGI endpoint they send via email). On the KBase endpoint, if you don’t see a directory with your name as in the example above, add /username to the Path for the KBase endpoint (/user/ is used in the example above).

You are then be able to drag data files or folders from the left endpoint to the right one, inside the Globus interface, and the files will appear in your KBase Staging Area.

Your transfer request will be submitted to Globus. If there are network or other problems, Globus will retry several times before giving up. Globus will send an email when the transfer is complete. In the KBase Staging Area, you may need to click on the refresh icon (two encircling arrows to the left of the user name) to update the *Import* tab.

<img src="/files/-M7tR7yKcefXYqOghkct" alt="" width="375">

### **Firewalls**

Some government institutions have a firewall between users and the outside world. In an effort to increase security, many web links can be blocked, including Globus traffic. If you see a message that ends with “Details: 500 Command failed….”, a firewall may be the cause.

### No ACL Rules

Another common error seen when using Globus shows and "No ACL rules" error on the KBase Bulk Share endpoint in the Globus UI. The most common cause for this error is multiple Globus accounts where the wrong account is linked to KBase. Be sure to log out of all Globus accounts, then link the correct account at <https://narrative.kbase.us/#/auth2/account>. If you still get this error, contact the help desk.&#x20;


# Transfer Data from JGI

The Data Transfer Service (DTS) connects KBase and the Joint Genome Institute (JGI) to seamlessly import data from JGI into your Narratives.

{% hint style="danger" %}

### **The first version of the DTS has been implemented in a beta state and there may be bugs. If you have trouble with transfers, please contact us at <engage@kbase.us>.**

Currently, the DTS only supports **isolate genomes which were sequenced by JGI**.
{% endhint %}

### **Before transferring data from JGI to KBase, you must:**

* [Get a JGI account](https://signon.jgi.doe.gov/signon) and link it to ORCID
* [Get a KBase account](http://kbase.us/sign-up-for-a-kbase-account/) and link it to ORCID (see the section account linking instructions [here](/manage-account/link-accounts) for help)
* [Use either Chrome, Firefox, or Safari browser](/getting-started/browsers)

{% embed url="<https://youtu.be/4CBBV-GlLB8?si=_zLCrNkx2Wi096x9>" %}
Tutorial on Data Transfer Service
{% endembed %}

### Selecting data

JGI lets users select data from the [JGI Genome Portal's](https://img.jgi.doe.gov/) sequencing projects and make them transferable to KBase, all in a few simple steps.&#x20;

The DTS operates as a push from IMG to KBase. You can find genome data to move to KBase using the normal IMG/MER Data Portal. Once you have a selction of genomes to push, add them to your cart.

{% hint style="info" %}
For more on finding JGI data, see their [series of webinar tutorials](https://mgm.jgi.doe.gov/img-webinar-series/).
{% endhint %}

### Transferring Data from JGI to KBase

When you have selected your genome(s) of interest from IMG, you can access the data through the Genome Cart.&#x20;

<figure><img src="/files/V6MxYWCOfcFQxeZvzGQp" alt=""><figcaption></figcaption></figure>

Select the genomes you want to transfer to KBase using the checkboxes in the IMG Genome Cart. When you have selected the genomes, go to the "Upload & Export & Save" tab.&#x20;

<figure><img src="/files/X43emTIRlC8yjDvRxSmS" alt=""><figcaption></figcaption></figure>

Provide a name and description for the transfer job, then click "Push to KBase" and it will begin the transfer.&#x20;

{% hint style="warning" %}
Some JGI data is on tape storage and not immediately accessible. If you choose archived genomes there may be a significant delay before the files are transferred.
{% endhint %}

When the transfer is complete you will get a confirmation that the files have successfully transferred to KBase Staging Area.&#x20;

<figure><img src="/files/WiMMMn5Ok9LigCpcM905" alt=""><figcaption></figcaption></figure>

To view the status of your transfer jobs, you can either click the link to "job status" in the confirmation, or go to "My Jobs" under the "My IMG" dropdown.

<figure><img src="/files/EMSawH6kWqburRvoZP1i" alt=""><figcaption></figcaption></figure>

### Importing to KBase

Open the Narrative where you want to import the data. Each transfer job that successfully moves files into KBase will have a folder with a name starting with `dts-`. Navigate into the folder you want to import.&#x20;

<figure><img src="/files/5mDObVa3H3jaRHjxonXF" alt=""><figcaption></figcaption></figure>

The top level will contain a file `manifest.json` and a folder `img`, if the transfer was successful. You can examine the contents of `img` to verify you have all the files you wish to import. The file `manifest.json` is formatted to be usable by the staging service for bulk import. Choose the type "Data Transfer Service Manifest" from the "Import As" dropdown, ensure the file is checked, and click "Import Selected."

<figure><img src="/files/8GWrVUOClRytQx10nq1Y" alt=""><figcaption></figcaption></figure>

The staging service will read the manifest and create a bulk import cell. From here, imports and data management proceed as normally. See [Importing Data](/data/upload-download-guide/uploads#bulk-import-guide) for more information.


# Using Apps

All about analysis tools available in KBase

1. How to access and run [analysis Apps in KBase](/apps/analysis)&#x20;
2. How to access [Beta Apps](/apps/beta) in KBase
3. Visit the[ KBase App Catalog](https://kbase.us/applist/)


# Analysis Apps in KBase

**KBase Apps** are analysis tools that you can use in KBase. Apps interoperate seamlessly to enable a range of scientific workflows (see figure below). Some of the apps were written by KBase scientists and developers; others are third-party tools that were integrated into KBase with our [Software Developer Kit (SDK)](/development/kbase-sdk). The number of apps available in KBase increases as members of the community use our SDK to integrate their analysis tools into the KBase platform.

### The App Catalog lists currently available Apps.&#x20;

* External [App Catalog](https://kbase.us/applist/) (no user account or sign-in required for browsing).
* Within KBase by clicking on the Catalog icon in the Menu <img src="/files/-MQJ2nifxGMUwo0HvzlT" alt="" data-size="line">&#x20;
* In a Narrative by clicking the right arrow at the top of the Apps Panel.

The majority of KBase Apps fall into the following categories:

* [Assembly](/apps/analysis/assembly-and-annotation#assembly)
* [Annotation](/apps/analysis/assembly-and-annotation#annotation)
* [Sequence Analysis](https://kbase.us/applist/#Sequence%20Analysis)
* [Comparative Genomics](/apps/analysis/comparative-genomics)
* [Metabolic Modeling](/apps/analysis/metabolic-modeling)
* [Expression](/apps/analysis/expression)
* [Microbial Communities](https://kbase.us/applist/#Microbial%20Communities)

{% hint style="info" %}
**Show me!**\
Go straight to the [App Catalog](https://narrative.kbase.us/#appcatalog)\
[![Screen Shot 2016-02-24 at 2.26.20 PM](/files/-LvwPpDm6hu34-m3fOwq)](https://narrative.kbase.us/#appcatalog)

Note: you will need a [KBase user account](/getting-started/sign-up#signing-up) to use our tools.
{% endhint %}

### Tips on using the [App Catalog](https://kbase.us/applist)

* Each app links to a reference page (which includes technical details about the inputs and outputs) called an App Details Page.
* To run apps, you will need to [sign in to the Narrative Interface](/getting-started/sign-up#signing-in).
* You can access the App Catalog from inside the Narrative Interface by clicking the small arrow in the upper right corner of the Apps Panel.

<img src="/files/-M79y2h-4bPDzoUCraai" alt="" width="563">

* You can click the star at the lower left of any app to add it to your “favorites.” The gray star will turn yellow to indicate that you have favorited the app. The number to the right of the star shows how many people have favorited that app.

<img src="/files/-M7A1MhGXhmT4e7jP_i3" alt="" width="375">

By default, the apps are sorted by category. Try the options in the “Organize by” menu to sort the apps by My Favorites, Run Count and other options.

![](/files/-LvwRqgIzLXUv9o5F-dZ)

### **Running Analysis Apps**

Once you have added an app to the Narrative, you will need to select a data object as the Input Object and set Parameters. When working within an application, a red bar or banner will highlight empty fields and errors.&#x20;

<div align="center"><img src="/files/-MgvrjKlZWB93ccOe8wh" alt="" width="563"></div>

### **More info**

For more information about using the App Catalog, see the [Narrative Interface User Guide](/getting-started/narrative).


# Assembly & Annotation

Some of the tools in KBase available for Assembly and Annotation

KBase provides multiple Apps for *de novo* [assembly](https://kbase.us/applist/#Genome%20Assembly) of prokaryotic Next-Generation Sequencing (NGS) reads from various sequencing platforms. These assemblies can then be [annotated](https://kbase.us/applist/#Genome%20Annotation) to explore structural and functional features of a Genome or use it in other analyses. The [interactive tutorials](/workflows/assembly-annotation) are a good way to learn about these workflows

### **Read Processing**

* [Trim Reads with Trimmomatic](https://kbase.us/applist/apps/kb_trimmomatic/run_trimmomatic/release) – Read trimming and adaptor removal
* [Filter Out Low-Complexity Reads with PRINSEQ](https://kbase.us/applist/apps/kb_PRINSEQ/execReadLibraryPRINSEQ/release) – Filter low complexity reads, i.e. polyG tails
* [Assess Read Quality with FastQC](https://kbase.us/applist/apps/kb_fastqc/runFastQC/release) – Quality assessment and reporting
* [Cutadapt](https://kbase.us/applist/apps/kb_cutadapt/remove_adapters/release) – Custom adapter removal at 3' and 5' ends
* [Filter Reads with Filtlong](https://kbase.us/applist/apps/kb_filtlong/run_kb_filtlong/release) – Filter long-read seqeuences

### Assembly

*De novo* [assembly](https://kbase.us/applist/#Genome%20Assembly) of Illumina and Ion Torrent next-generation sequencing reads. Supports single-end and paired-end read libraries.

* [Assemble with HipMer](https://kbase.us/applist/apps/hipmer/run_hipmer_hpc/release) – [HipMer](https://sourceforge.net/p/hipmer/wiki/Home/) is a highly-parallelized port of JGI’s Meraculous assembler. Meraculous is a de Bruijn graph-based which increases speed by not performing error correction. Instead, it bases contigs on already high-quality scores and fills the gaps based on localized assemblies from the reads. HipMer enhances the speed of Meraculous.
* [Assemble with IDBA-UD](https://kbase.us/applist/apps/kb_IDBA/run_idba_ud/release) – [IDBA-UD](http://i.cs.hku.hk/~alse/hkubrg/projects/idba_ud/) is an iterative graph-based assembler for single-cell and standard short read data and is good for data of highly uneven sequencing depth. This assembler uses an iterative approach for selecting k-mer size that compensates for the information loss associated with single k-mer based de Bruijn graphs, making IDBA-UD one of the more accurate microbial assemblers.
* [Assemble with MaSuRCA](https://kbase.us/applist/apps/kb_MaSuRCA/run_masurca_assembler/release) – [MaSuRCA](https://academic.oup.com/bioinformatics/article/29/21/2669/195975/The-MaSuRCA-genome-assembler) is a short read assembler that combines the benefits of de Bruijn graph and overlap layout consensus assembly approaches. The main concept is the creation of super-reads that contain sequence information present in the original reads, which super-reads are then extended in both directions using an efficient k-mer lookup table. MaSuRCA is one of a smaller set of assemblers biologists use for eukaryotic assembly.
* [Assemble with MEGAHIT](https://kbase.us/applist/apps/MEGAHIT/run_megahit/release) – [MEGAHIT](https://academic.oup.com/bioinformatics/article-lookup/doi/10.1093/bioinformatics/btv033) is a single node assembler for large and complex metagenomics NGS reads. It makes use of succinct de Bruijn graph (SdBG) to achieve low memory assembly, making it fast and especially suitable for assembly of small metagenomes, metatranscriptomes or low-coverage data in general.
* [Assemble with SPAdes](https://kbase.us/applist/apps/kb_SPAdes/run_SPAdes/release) – [SPAdes](http://online.liebertpub.com/doi/full/10.1089/cmb.2012.0021) is a single-cell and standard assembler based on paired de Bruijn graphs, considered to be one of the most accurate microbial assemblers. SPAdes employs a multisized de Bruijn graph which detects and removes bubble and chimeric reads, estimates insert distance from paired kmers, and computes contigs based on paired assembly graph.
* [Assemble with Velvet](https://kbase.us/applist/apps/Velvet/run_velvet/release) – [Velvet](http://onlinelibrary.wiley.com/doi/10.1002/0471250953.bi1105s31/full) is a classic de Bruijn graph based assembler that works by efficiently manipulating de Bruijn graphs through simplification and compression. It eliminates errors and resolves repeats by first using an error correction algorithm that merges sequences together. Repeats are then removed from the sequence via the repeat solver that separates paths which share local overlaps.
* [Assemble Long Reads with Flye](https://kbase.us/applist/apps/kb_flye/run_kb_flye/release) – [Flye](https://www.nature.com/articles/s41587-019-0072-8) is a de novo assembler for long-read, i.e. single-molecule sequencing reads, such as those produced by PacBio and Oxford Nanopore Technologies.
* [Assemble Reads with Unicycler](https://kbase.us/applist/apps/kb_unicycler/run_unicycler/release) – [Unicycler](http://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1005595) is an assembly pipeline for bacterial genomes that can perform hybrid assemblies with both Illumina reads and PacBio or Nanopore through SPAdes-optimized and miniasm+Racon. It can also run solely short- or long-read assemblies.&#x20;
* [Compare assemblies with QUAST](https://kbase.us/applist/apps/kb_quast/run_QUAST_app/release) – Assess the output assemblies from different configurations of the same assembler, or compare assemblies from multiple assemblers to determine which one is optimal for downstream analysis.
* [Polish Assemblies with Polypolish](https://kbase.us/applist/apps/kb_polypolish/run_kb_polypolish/release) – [Polypolish](https://doi.org/10.1371/journal.pcbi.1009802) is a tool for polishing genome assemblies with short reads.

### Annotation

Genomes can be [annotated](https://kbase.us/applist/#Genome%20Annotation) with Prokka or RAST.&#x20;

* [Annotate Domains in a Genome](https://kbase.us/applist/apps/DomainAnnotation/annotate_domains_in_a_genome/release) – identifies protein domains from widely used domain libraries (COGs, TIGRfams, Pfam).
* [Annotate Assembly with Prokka](https://kbase.us/applist/apps/ProkkaAnnotation/annotate_contigs/release) – combines multiple open-source annotation tools in a quick and thorough annotation pipeline for prokaryotic sequences for genomes, plasmids, and metagenomes.
* [Annotate Microbial Assembly](https://kbase.us/applist/apps/RAST_SDK/annotate_contigset/release) – uses components from the RAST ([Rapid Annotations using Subsystems Technology](http://rast.nmpdr.org/)) toolkit to annotate an assembled bacterial or archaeal genome.
* [Annotate Microbial Genome](https://kbase.us/applist/apps/RAST_SDK/reannotate_microbial_genome/release) – uses RAST to annotate a prokaryotic genome, to update the annotations of a genome, or to perform computations on a set of genomes so that they are consistent.
* [Annotate Plant Coding Sequences with Metabolic Functions](https://kbase.us/applist/apps/kb_plant_rast/annotate_plant_transcripts/release) – performs functional annotation of plant cDNA or protein sequences.
* [Bulk Annotate Genomes/Assemblies](https://kbase.us/applist/apps/RAST_SDK/bulk_annotate_genomes_assemblies/release) – uses components from the RAST ([Rapid Annotations using Subsystems Technology](http://rast.nmpdr.org/)) toolkit to annotate a set of genomes or assemblies.&#x20;
* Annotate and Distill [Assemblies](https://kbase.us/applist/apps/kb_DRAM/run_kb_dram_annotate/release) / [Genomes](https://kbase.us/applist/apps/kb_DRAM/run_kb_dram_annotate_genome/release) with DRAM  – [DRAM](https://academic.oup.com/nar/article/48/16/8883/5884738) (Destilled and Refined Annotation of Metabolism) provides metabolic summaries by predicting coding sequences and annotations microbial sequences from assemblies, isolates, and metagenome-assembled genomes.&#x20;
* [Annotate with Snekmer Apply](https://kbase.us/applist/apps/SnekmerLearnApply/run_SnekmerLearnApply/release) – [Snekmer Apply](https://pubmed.ncbi.nlm.nih.gov/36789294/) will annotate genomes or protein seqeuences to provide updated protein sequences and gene ontologies.&#x20;

*The output of the annotation apps is a Genome, which is displayed in a tabular genome viewer (see below) that shows information about the Genome as well as a list of contigs and the genes that were called on each contig.*

![ViewContig](/files/-LvwVVOjY_j9H856wEXV)


# Comparative Genomics

Some of the tools in KBase available for comparative genomics and phylogenetics

KBase provides multiple [comparative genomics and phylogenetic analysis Apps](https://kbase.us/applist/#Comparative%20Genomics) to enable researchers to understand evolutionary relationships between organisms and explore structural and functional variance across genomes.

### **Phylogenetic and Taxonomic Analysis**

* [Insert Genome Into Species Tree](https://kbase.us/applist/apps/SpeciesTreeBuilder/insert_set_of_genomes_into_species_tree/release) – Build a phylogenetic tree with your genomes and RefSeq microbial genomes based on conserved orthologous groups
* [Build Microbial Species Tree](https://narrative.kbase.us/#catalog/apps/kb_phylogenomics/build_microbial_speciestree/beta) ([in beta](/apps/beta)) – Build a phylogenetic tree with your genome and phylum exemplars from RefSeq
* [Build Gene Tree](https://narrative.kbase.us/#catalog/apps/kb_phylogenomics/build_gene_tree/beta) ([in beta](/apps/beta)) – Build an evolutionary reconstruction tree for a collection of sequence homology related genes
* [Build Phylogenetic Tree with FastTree2](https://kbase.us/applist/apps/kb_fasttree/run_FastTree/release)  – Build a tree with FastTree2 directly on a user-created Multiple Sequence Alignment (MSA)

### Sequence Homology and Functional Analysis

* [BLASTx](https://kbase.us/applist/apps/kb_blast/BLASTx_Search/release), [tBLASTn](https://kbase.us/applist/apps/kb_blast/tBLASTn_Search/release), and [tBLASTx](https://kbase.us/applist/apps/kb_blast/tBLASTx_Search/release) – Search genes in genomes and Annotated Metagenome Assemblies (AMAs) and compare translated protein sequences
* [BLASTn](https://kbase.us/applist/apps/kb_blast/BLASTn_Search/release) – Search genes and AMAs and compare nucleotide sequences
* [BLASTp](https://kbase.us/applist/apps/kb_blast/BLASTp_Search/release) and [psiBLAST](https://kbase.us/applist/apps/kb_blast/psiBLAST_msa_start_Search/release) – Search genes and AMAs and compare protein sequences
* MUSCLE for [nucleotide](https://kbase.us/applist/apps/kb_muscle/MUSCLE_nuc/release) or [protein](https://kbase.us/applist/apps/kb_muscle/MUSCLE_prot/release) sequences – Build and analyze multiple sequence alignments (MSA)
* [HMMER MSA Collection](https://narrative.kbase.us/#catalog/modules/kb_hmmer) – Search and profile genes in genomes and Annotated Metagenome Assemblies (AMAs) with a collection of user-defined MSAs
  * [HMMER Search from MSA](https://kbase.us/applist/apps/kb_hmmer/HMMER_MSA_Search/release) – Search genes in genomes and AMAs with a single user-defined MSA &#x20;
  * [HMMER Search with dbCAN2 of CAZy families](https://kbase.us/applist/apps/kb_hmmer/HMMER_dbCAN_Search/release) – Search and profile genes in genomes and AMAs with the suite of CAZy dbCAN2 hidden markov models (HMMs)
  * [HMMER Search for Phylo Markers](https://kbase.us/applist/apps/kb_hmmer/HMMER_PhyloMarkers_Search/release) – Search and profile genes in genomes and AMAs with the suite of GTDB protein phylogenetic marker HMMs
  * [HMMER Search for Env BioElement families](https://kbase.us/applist/apps/kb_hmmer/HMMER_env-bioelement-hmm_Search/release) – Search and profile genes in genomes and AMAs with the suite of Environmental BioElement metabolic families
* [Annotate Domains in a GenomeSet](https://kbase.us/applist/apps/kb_phylogenomics/run_DomainAnnotation_Sets/release) – Compute canonical gene family membership using COG, Pfam, and TIGRFAMs&#x20;
* [View Function Profile for Genomes](https://kbase.us/applist/apps/kb_phylogenomics/view_fxn_profile/release) and [View Function Profile for a Phylogenetic Tree](https://kbase.us/applist/apps/kb_phylogenomics/view_fxn_profile_phylo/release) – Examine functional content from canonical domain family annotation
* [Run Fama Genome Profile ](https://kbase.us/applist/apps/FamaProfiling/run_FamaGenomeProfiling/release)– Examine function and taxonomic profiling of microbiomes and functional genes through a similarity search for all predicted proteins in a genome using fast aligner DIAMOND and customized reference protein databases.&#x20;

### **Pangenome Exploration**&#x20;

* [Compare Two Proteomes](https://kbase.us/applist/apps/GenomeProteomeComparison/compare_two_proteomes/release) – Compare proteomes and produce a synteny map dot plot matrix and table of gene differences
* [Build Pangenome with OrthoMCL](https://kbase.us/applist/apps/PangenomeOrthomcl/build_pangenome_with_orthomcl/release) – Group orthologous protein sequences and identify core and non-core genes across multiple species
* [Calculate Pangenome wiht mOTUpan](https://narrative.kbase.us/#catalog/apps/kb_motupan/run_kb_motupan/release) - Calculate pangenome for microbial genomes, including MAGs of varying quality
* [Compute Pangenome](https://kbase.us/applist/apps/GenomeComparisonSDK/build_pangenome/release) – Build a pangenome and evaluate protein family conservation&#x20;
* [Pangenome Circle Plot ](https://kbase.us/applist/apps/kb_phylogenomics/view_pan_circle_plot/release)– View a microbial pangenome as a circle plot using &#x20;
* [Compare Genomes from Pangenome](https://kbase.us/applist/apps/GenomeComparisonSDK/compare_genomes/release) – Compare isofunctional and homologous gene families within a pangenome
* [Phylogenetic Pangenome Accumulation](https://kbase.us/applist/apps/kb_phylogenomics/view_pan_phylo/release) – View the pangenome in a phylogenetic context&#x20;


# Metabolic Modeling

Some of the tools in KBase available for metabolic modeling

KBase has a suite of [Apps supporting the reconstruction, prediction, and design of metabolic models](https://kbase.us/applist/#Metabolic%20Modeling) in microbes and plants. Genome-scale metabolic models can be used to explore an organism’s growth in specific media conditions, determine which biochemical pathways are present, optimize production of an important metabolite, identify high flux pathways, and more.

### **Flux Balance Analysis**

* [MS2 - Build Prokaryotic Metabolic Models with OMEGGA](https://kbase.us/applist/apps/ModelSEEDReconstruction/build_metabolic_models/release)  – Construct genome-scale metabolic models based on a bacterial genome and media conditions utilizing omics-enabled global gapfilling (OMEGGA)&#x20;
  * Red - consumed, green - produced, blue - not consumed
* [MS2 - Improved Gapfill Metabolic Model with OMEGGA](https://kbase.us/applist/apps/ModelSEEDReconstruction/gapfill_metabolic_models/release) – Fill in missing reactions based on stoichiometry
* [Run Flux Balance Analysis](https://kbase.us/applist/apps/fba_tools/run_flux_balance_analysis/release) – Predict metabolic fluxes based on steady state analyses
* [Compare FBA solutions](https://kbase.us/applist/apps/fba_tools/compare_fba_solutions/release) – Determine optimal conditions of flux&#x20;
* [Check Model Mass Balance](https://kbase.us/applist/apps/fba_tools/check_model_mass_balance/release) – Ensure accuracy&#x20;
* [Compare Models](https://kbase.us/applist/apps/fba_tools/compare_models/release) – View multiple models side by side

### Editing Models

* [Edit Metabolic Model](https://kbase.us/applist/apps/fba_tools/edit_metabolic_model/release) – Build unique models suited to specific experiments.&#x20;
* [Create or Edit Media](https://kbase.us/applist/apps/fba_tools/edit_media/release) – Create specialized growth conditions
* [Bulk Download Modeling Objects](https://kbase.us/applist/apps/fba_tools/bulk_download_modeling_objects/release) – Save modeling data for future analysis&#x20;

### Comparative Genomics

* [Propagate Model to New Genome](https://kbase.us/applist/apps/fba_tools/propagate_model_to_new_genome/release) – Translate metabolic models from one organism to another

### Expression

* [Compare Flux with Expression ](https://kbase.us/applist/apps/fba_tools/compare_flux_with_expression/release)– Compare reaction fluxes with gene expression values to identify metabolic pathways where expression and flux data agree or conflict and compare with chemical abundance data
* [Simulate Growth on Phenotype Data](https://kbase.us/applist/apps/fba_tools/simulate_growth_on_phenotype_data/release) – Reconcile models with empirical data

### **Metabolomics**

* [Escher Pathway Viewer App](https://narrative.kbase.us/#catalog/apps/kb_escher/run_kb_pathway_view/beta) ([in beta](/apps/beta)) – Display Escher metabolic pathways and combine pathways with flux balance analysis and metabolomics data to view flux or expression.&#x20;
  * When viewing flux is in green and expression is in brown. The oval size and color intensity reflects metabolite abundance.&#x20;
* [PickAxe App](<https://narrative.kbase.us/#catalog/apps/kb_pickaxe/pickaxe/beta >) ([in beta](/apps/beta)) – Generate novel compounds from an FBA Model and a Chemical Abundance Matrix&#x20;
  * Pickaxe is part of the MINE Databases, documentation for which can be found here: <https://github.com/tyo-nu/MINE-Database/blob/master/minedatabase/pickaxe.py>.&#x20;
* [Fit Model to Exometabolite Data](https://narrative.kbase.us/#catalog/apps/fba_tools/fit_exometabolite_data/beta) ([in beta](/apps/beta)) – Identify biochemical reactions to add to a draft metabolic model for production and consumption of exometabolites&#x20;

### **Microbial Communities**

* [Merge Metabolic Models into Community Model](https://kbase.us/applist/apps/fba_tools/merge_metabolic_models_into_community_model/release) – Investigate community metabolism&#x20;
* [Build Metagenome Metabolic Model](https://kbase.us/applist/apps/fba_tools/build_metagenome_model/release) – Build a metagenome metabolic model from an annotated assembly or bins


# Metagenomics & Community Exploration

Some of the tools in KBase available for analyzing microbial communities and metagenomics

KBase has a multitude of [Apps designed to analyze microbial communities and metagenomic reads data](https://kbase.us/applist/#Microbial%20Communities). They provide a way to understand the functional interactions between species in microbial communities. Annotated genomes extracted from metagenomic sequence libraries can be used for metabolic modeling, comparative phylogenomics, functional profiling, and more.

### **Metagenomic Read Analysis & Assembly**

* [FastQC](https://kbase.us/applist/apps/kb_fastqc/runFastQC/release)  – Raw read quality control and analysis&#x20;
* [Trim Reads with Trimmomatic](https://kbase.us/applist/apps/kb_trimmomatic/run_trimmomatic/release) – Read trimming and adaptor removal
* [Kaiju](https://kbase.us/applist/apps/kb_kaiju/run_kaiju/release) – Taxonomy assignment of raw metagenomic reads
* [MEGAHIT](https://kbase.us/applist/apps/MEGAHIT/run_megahit/release) – Assemble large and complex metagenomic reads
* [MetaSPAdes](https://kbase.us/applist/apps/kb_SPAdes/run_metaSPAdes/release) – Assemble shotgun metagenomic reads
* [IDBA-UD](https://kbase.us/applist/apps/kb_IDBA/run_idba_ud/release) – Assemble short metagenomic reads

### Contig Binning, Optimization, and Quality Analysis&#x20;

* [MaxBin2](https://kbase.us/applist/apps/kb_maxbin/run_maxbin2/release) – Group metagenomic contigs using depth-of-coverage, nucleotide composition, and marker genes
* [MetaBAT2](https://kbase.us/applist/apps/metabat/run_metabat/release) – Group metagenomic contigs using strain abundance and nucleotide composition
* [CONCOCT](https://kbase.us/applist/apps/kb_concoct/run_kb_concoct/release) – Group metagenomic contigs using depth-of-coverage and nucleotide composition
* [DAS-Tool](https://kbase.us/applist/apps/kb_das_tool/run_kb_das_tool/release) – Optimizes bacterial or archaeal genome bins using a de-replication, aggregation and scoring strategy
* [CheckM ](https://kbase.us/applist/apps/kb_Msuite/run_checkM_lineage_wf/release)– Assess genome and MAG quality
* [Jorg](https://kbase.us/applist/apps/kb_jorg/run_kb_jorg/beta) ([in beta](/apps/beta))– Improve and circularize single-genome assemblies and MAGs
* [Circos](https://kbase.us/applist/apps/kb_circos/run_kb_circos/beta) ([in beta](/apps/beta)) – Visualize assembly coverage&#x20;

### **Community Exploration**

* RAST for [Multiple Microbial Genomes](https://kbase.us/applist/apps/RAST_SDK/reannotate_microbial_genomes/release) and [Multiple Microbial Assemblies](https://kbase.us/applist/apps/RAST_SDK/annotate_contigsets/release) – Annotate multiple bacterial or archaeal genomes and assemblies or sets using RASTtk
* [Annotate and Distill Genomes with DRAM](https://kbase.us/applist/apps/kb_DRAM/run_kb_dram_annotate/release) – Annotates Metagenome Assembled Genomes (MAGs) and provides an interactive functional summary per genome
* [dRep](https://kbase.us/applist/apps/kb_dRep/run_dereplicate/beta) ([in beta](/apps/beta)) – Dereplicate binned contigs, assemblies, and genomes using nucleotide similarity&#x20;
* [VirSorter](https://kbase.us/applist/apps/VirSorter/run_VirSorter/release) – Identify viral sequences from viral and microbial metagenomes
* [Classify Microbes with GTDB-Tk](https://kbase.us/applist/apps/kb_gtdbtk/run_kb_gtdbtk_classify_wf/release)  – Assign taxonomy to isolate and classify Metagenome Assembled Genomes (MAGs) using Genome Taxonomy Database (GTDB) genome taxonomy
* [ModelSEED](https://kbase.us/applist/apps/fba_tools/build_multiple_metabolic_models/release) – Generate and merge metabolic models from annotated genomes
* [Phylogenetic Pangenome Accumulation](https://kbase.us/applist/apps/kb_phylogenomics/view_pan_phylo/release) – Produce phylogenetic trees from annotated genomes
* [Metabolic modeling tools for microbial communities](/apps/analysis/metabolic-modeling#microbial-communities) – analyze community metabolism


# Data Matrices - Amplicon, Stats

Some of the tools in KBase available for rarefaction, standardization, and analyzing amplicon matrices and OTU tables.

### Data Cleaning

Rarefaction and standardization can be performed in KBase to track data provenance on how the data was transformed. These Apps can be used in any order and/or repeatedly. For instance to remove singletons, rarefy samples, and then calculate relative abundance of rarefied samples.&#x20;

* [Transform Matrix App](https://kbase.us/applist/apps/GenericsAPI/transform_matrix/release) – uses an AmpliconMatrix object as input
  * provides abundance filtering, which can remove rows or columns if all values fall below a user-specified threshold and/or if the sum of all values falls below a user-specified threshold
  * perform standardization, with options for centering the data before scaling, scaling the data to unit variance, and whether to perform the standardization on columns or rows
  * perform ratio transformation, by either Centre or Isometric Log Ratio Transformation methods, and be applied to columns or rows&#x20;
* [Rarefy Matrix App](https://kbase.us/applist/apps/GenericsAPI/rarefy_matrix/release) – allows manual selection of a seed value to aid reproducibility to columns or rows

### Taxonomic Classification

* Classify rRNA with taxonomy using naïve Bayes with RDP Classifier App (in dev/beta) – assigns taxonomies to amplicon matrices based on several gene databases including RDP Classifier 2.13 - 16srna, RDP Classifier - fungallsu, and SILVA 138 - SSU, Full Length. *Silva databases are currently under development*.
* [Taxonomy Abundance Barplot App](https://narrative.kbase.us/#catalog/apps/TaxonomyAbundance/run_TaxonomyAbundance/beta) (in beta) – visualize the taxonomic abundance

### Statistical Analysis

Statistical comparisons can be calculated, such as clustering (hierarchical or k-means) and PCA analysis.&#x20;

* [Perform NMDS Analysis App](https://narrative.kbase.us/#catalog/apps/kb_Amplicon/run_metaMDS/beta) ([in beta](/apps/beta)) – performs Non-metric Multidimensional Scaling Analysis on amplicon matrix data to visualize it in two dimensions
* [Perform Categorical Variable Statistics Analysis App](https://kbase.us/applist/apps/GenericsAPI/perform_variable_stats/release) – perform Categorical Variable Statistics Analysis on an amplicon matrix object
* [Perform Similarity Percentage (SIMPER) Statistics Analysis App](https://kbase.us/applist/apps/GenericsAPI/perform_simper/release) – performs similarity percentage or SIMPER (Clarke 1993) analysis based on the decomposition of Bray-Curtis dissimilarity index

### Functional Prediction for Amplicon Matrices

* [Map Tax to Functions using FAPROTAX App](https://narrative.kbase.us/#catalog/apps/kb_faprotax/faprotax/beta) – identify functions and assign them to a new AmpliconMatrix using a manually-curated database that focuses on marine and lake biochemistry, such as sulfur, nitrogen, hydrogen, and carbon cycling.

{% hint style="info" %}
[Louca Lab FAPROTAX site](http://www.loucalab.com/archive/FAPROTAX/lib/php/index.php?section=Home) (<http://www.loucalab.com/archive/FAPROTAX/>)

Louca, S., Parfrey, L.W., Doebeli, M. (2016) Decoupling function and taxonomy in the global ocean microbiome. Science **353**: 1272-1277. <https://www.science.org/doi/10.1126/science.aaf4507>
{% endhint %}

* [Predict prokaryote EC, KO, & MetaCyc function abundances with PICRUSt2 App](https://narrative.kbase.us/#catalog/apps/kb_PICRUSt2/run_picrust2_pipeline/beta) – identify functions from many different ontologies, including COG, EC, and KO, or any combination of which can be chosen in the parameters. Create new FunctionalProfile objects for the SampleSet, AmpliconMatrix, or both, and outputs a new AmpliconMatrix Object.&#x20;

{% hint style="info" %}
Douglas, G.M., Maffei, V.J., Zaneveld, J.R. *et al.* (2020). PICRUSt2 for prediction of metagenome functions. *Nat Biotechnol* **38:** 685–688 (2020). <https://doi.org/10.1038/s41587-020-0548-6>

[PICRUSt2: An improved and extensible approach for metagenome inference](https://www.biorxiv.org/content/10.1101/672295v1.full.pdf) (preprint)
{% endhint %}

Most of the relevant KBase Apps can be found in the Utilities section of the Apps panel and filtered through searching with "GenericsAPI" or the [GenericsAPI Module Catalog](https://narrative.kbase.us/#catalog/modules/GenericsAPI).&#x20;


# Chemical Abundance

Use abundance matrices with element, compound, or exometabolite data for analysis.

## Data Cleaning and Statistical Analysis

As both chemical abundance and amplicon matrices are similar, similar apps for cleaning, normalizing, and performing statistical comparisons can be done on both types. See the previous page on [Data Matrices](/apps/analysis/matrix) for information on using these apps.

## Popular Use Cases of Chemical Abundance Data

You can view a few public Narratives that demonstrate popular use cases for chemical abundance tables in KBase.

* Metabolomics (and other omics) data representation in Escher pathway maps
  * <https://narrative.kbase.us/narrative/55494>
  * <https://narrative.kbase.us/narrative/86637>&#x20;
* Cheminformatics expansion of known metabolites (based on known compounds in a metabolic model) and mapping onto unknown peaks from metabolomics data (PNNL Summer School workflows - WHONDRS)
  * <https://narrative.kbase.us/narrative/61753>
* Reconcile metabolic models to exo-metabolite data - Web of Microbes
  * <https://narrative.kbase.us/narrative/39597>


# Expression & Transcriptomics

Some of the tools in KBase available for expression analysis

KBase offers a powerful suite of [expression analysis Apps](https://kbase.us/applist/#Expression). Starting with short reads, you can use the tool suite to analyze transcriptomic data in a . You can also compare the expression data with the flux when studying metabolic models in KBase and identify pathways where expression and flux agree or conflict.

KBase offers a powerful suite of [expression analysis tools](https://kbase.us/applist/#Expression). Starting with short reads, you can use the tool suite to assemble, quantify long transcripts, and identify differentially expressed genes. You can also compare the expression data with the flux when studying metabolic models in KBase and identify pathways where expression and flux agree or conflict

### **Reads Management**

* [Assess Read Quality with FastQC](https://kbase.us/applist/apps/kb_fastqc/runFastQC/release) – Read quality analysis
* [Cutadapt](https://kbase.us/applist/apps/kb_cutadapt/remove_adapters/release) – Remove adapter sequences from reads
* [Trim Reads with Trimmomatic](https://kbase.us/applist/apps/kb_trimmomatic/run_trimmomatic/release) – Read trimming and removing Illumina adapters

### **Alignment**

* [Align Reads using Bowtie2](https://kbase.us/applist/apps/kb_Bowtie2/align_reads_using_bowtie2/release) – aligns the sequencing reads for a set of two or more samples to long reference sequences of a prokaryotic genome using Bowtie2 and outputs a set of alignments for the given sample set in BAM format.
* [Align Reads using TopHat2](https://kbase.us/applist/apps/kb_tophat2/align_reads_using_tophat2/release) – aligns the sequencing reads for a set of two or more samples to an eukaryotic genome using TopHat2 in order to identify splice junctions between exons with the help of Bowtie2 mapping program.
* [Align Reads using HISAT2](https://kbase.us/applist/apps/kb_hisat2/align_reads_using_hisat2/release) – aligns the sequencing reads for a set of two or more samples to long reference sequences of a genome using HISAT2 and outputs a set of alignments for the given sample set or reads set in BAM format.
* [Align Reads using STAR](https://narrative.kbase.us/#catalog/apps/STAR/align_reads_using_STAR/beta) – aligns the sequencing reads of a single or a set of two (paired end) reads to long reference sequences of a prokaryotic genome using the STAR alignment program.

### **Assembly**

* [Assemble Transcripts using Cufflinks](https://kbase.us/applist/apps/kb_cufflinks/assemble_transcripts_using_cufflinks/release) – assembles transcripts for a given sample or a sample set using Cufflinks so that you can view the relative abundances of the assembled transcripts in a histogram, obtained as output from Cufflinks.
* [Assemble Transcripts using StringTie](https://kbase.us/applist/apps/kb_stringtie/run_stringtie/release) – assembles transcripts for a given sample or a sample set using StringTie so that you can view the relative abundances of the assembled transcripts in a histogram.

### **Differential Expression**&#x20;

* [Identify Differential Expression using Cuffdiff](https://kbase.us/applist/apps/kb_cufflinks/run_Cuffdiff/release) – uses the Cufflinks transcripts for two or more samples to calculate gene and transcript levels in more than one condition and finds significant changes in the expression levels.
* [Create Differential Expression Matrix using Ballgown](https://kbase.us/applist/apps/kb_ballgown/run_ballgown_app/release) – uses the transcripts for two or more samples obtained from either Cufflinks or StringTie to calculate gene and transcript levels in more than one condition and finds significant changes in the expression levels.
* [Create Differential Expression Matrix using DESeq2](https://kbase.us/applist/apps/kb_deseq/run_DESeq2/release) – uses the transcripts for two or more samples obtained from either Cufflinks or StringTie to calculate gene and transcript levels in more than one condition and finds significant changes in the expression levels.

### Downstream Analyses

* [Filter Expression Matrix](https://kbase.us/applist/apps/CoExpression/expression_toolkit_filter_expression/release) – Filter an expression matrix using either Log Odds Ratio (LOR) or ANalysis of VAriance (ANOVA) algorithms.
* [Cluster Expression Data – Hierarchical ](https://kbase.us/applist/apps/KBaseFeatureValues/expression_toolkit_cluster_hierarchical/release)– Perform hierarchical clustering to group gene expression data into a dendrogram.
* [Estimate K for K-means Clustering](https://kbase.us/applist/apps/KBaseFeatureValues/expression_toolkit_estimate_k/release) – Generate reasonable numbers of clusters (K) for use in the [Cluster Expression Data - K-Means](https://kbase.us/applist/apps/KBaseFeatureValues/expression_toolkit_cluster_k_means/release) App.
* [Cluster Expression Data – K-means](https://kbase.us/applist/apps/KBaseFeatureValues/expression_toolkit_cluster_k_means/release) – Perform K-means clustering to group expression data for observing and analyzing patterns of gene expression.
* [Cluster Expression Data – WGCNA](https://kbase.us/applist/apps/CoExpression/expression_toolkit_cluster_WGCNA/release) – Perform weighted gene co-expression network analysis (WGCNA) to detect gene clusters and expression patterns.
* [View Interactive Heatmap](https://kbase.us/applist/apps/NarrativeViewers/view_expression_interactive_heatmap/release) – display heatmap of expressed genes.&#x20;
* [View Multi-cluster Heatmap](https://kbase.us/applist/apps/CoExpression/expression_toolkit_view_heatmap/release) – display multi-cluster heatmap of expressed genes.


# Apps in Beta

These KBase Apps are being tested and debugged.

![App Panel Menu](/files/-M8MObUWmm9AbzxFDwDf)

To use apps still in development, you can click the “R” in the Apps Panel, that changes it to a “B” which shows apps that are in Beta. Click the "OK" button to start using Beta Apps.&#x20;

{% hint style="info" %}
Choosing Beta will only show the apps that are still being tested and debugged. Click the "B" icon to go back to Released apps.&#x20;
{% endhint %}

![Click the "R" icon in the Apps menu to toggle to "B" version of available apps. ](/files/-M8MCM0-qNbE-QN9xO9D)


# Running Common Workflows

User guides on how to run different workflows in KBase.

Here is an example outline of the major workflows and datatypes in KBase. Common workflows include genome assembly and annotation, comparing genomes through taxonomy, sequence analysis, metabolic modeling, transcriptomics and expression analysis, viral analysis, and community analysis through metagenomics.&#x20;

<figure><img src="/files/dh6pAZ0BoI6TOdQWIBjt" alt=""><figcaption><p>Examples of common analysis workflows in KBase</p></figcaption></figure>

Find example workflows here:

1. [Assembling & Annotating Microbial Genomes](/workflows/assembly-annotation)
2. [Comparative Genomics & Phylogenetic Analysis](/workflows/comparative-genomics)
3. [Constructing Metabolic Models](/workflows/metabolic-models)
4. [Metagenomic & Community Analysis](/workflows/metagenomic-analysis)
5. [RNAseq & Expression Analysis](/workflows/rnaseq)
6. [Samples & Metadata](/data/samples)


# Assembling & Annotating Microbial Genomes

Running assembly and annotation in KBase

KBase enables *de novo* assembly of prokaryotic NGS reads from various sequencing platforms. Assemblies can then be annotated with RAST or Prokka, enabling you to explore structural and functional features of a Genome or use it in other analyses.&#x20;

{% hint style="info" %}
Steps for [assembly and annotation](/apps/analysis/assembly-and-annotation) include:

1. Upload reads to assemble.
2. [Add your Reads data object](/data/upload-download-guide/reads) to the Narrative.
3. Search the App Catalog and insert an [Assembly App](https://kbase.us/applist/#Genome%20Assembly)&#x20;
   * Use one of these tools to create an Assembly object.
4. Search the App Catalog and insert an [Annotation App](https://kbase.us/applist/#Genome%20Annotation)
   * Use one of these tools to annotate the Assembly object.
5. Examine assembly job statistics and genome annotation.
   {% endhint %}

## Narrative Tutorials

* [Microbial Genomics in KBase: Drafting Isolate Genomes (Short-reads)](https://narrative.kbase.us/narrative/83666) – demonstrates how to draft isolate genomes using Illumina short-read sequencing from quality assessment to assembly and annotation.&#x20;
* [Microbial Genomics: Drafting Isolate Genomes (Long-reads)](https://narrative.kbase.us/narrative/200668) – demonstrates how to draft isolate genomes using Oxford Nanoport MinION long-read sequences from quality assessment to assembly and annotation.&#x20;

## Video Tutorials

{% embed url="<https://youtu.be/EGW3rA8tWf4>" %}
Introduction to Genome Analysis
{% endembed %}

{% embed url="<https://youtu.be/W6QBVmb9vC0?si=pA8HuO51SZfECG0d>" %}
Microbial Genomics in KBase (short- and long-read examples)
{% endembed %}


# FAQ: Assembly and Annotation

## Before submitting them for annotation. Is there a direct way to import to JGI or do we have to save FASTA files and separately submit to JGI/IMG?&#x20;

This page on transferring data from JGI to KBase takes you through the process. You can also start from the KBase search and use the JGI tab.&#x20;

## What is the typical threshold to determine whether our assembled genome is contaminated?&#x20;

In general, a genome that is >90% complete and <5% contaminated is high quality. A rough guide to the quality of MAGs and SAGs can be found here: <https://www.nature.com/articles/nbt.3893>&#x20;

## Is there a limit on the size of the FASTQ files I can upload to KBase?&#x20;

Assemblers currently have an upper limit of between 180,263,840 paired reads and 240,351,788 reads depending on complexity. If the job has been run twice, exceeded the 7 day limit, and your data is in this size range, it may be too big for KBase at this time.&#x20;

## How do I remove adapters?

Based on sequencing metadata and FastQC results, you can use available read processing tools, such as [Trimmomatic](https://kbase.us/applist/apps/kb_trimmomatic/run_trimmomatic/release) or [Cutadapt](https://kbase.us/applist/apps/kb_cutadapt/remove_adapters/release), to remove custom 3' and 5' adapters. &#x20;

## What tools are available for removing polyG tails?

If the reads have polyG tails, you can filter out the low-complexity reads using [PRINSEQ](https://kbase.us/applist/apps/kb_PRINSEQ/execReadLibraryPRINSEQ/release). [Filtlong](https://kbase.us/applist/apps/kb_filtlong/run_kb_filtlong/release) can also be used for long-read sequencing libraries to filter out low-complexity reads.&#x20;

## What is the best assembler?&#x20;

The “best” assembler often depends on the user. Sometimes the user may want the most contiguous metagenome, or an assembly with minimized assembly artifacts. Users can use multiple assemblers and choose whichever results in the best assembly for their purpose.&#x20;

## How do I compare the results of various assemblers?&#x20;

You can use the [QUAST App](https://kbase.us/applist/apps/kb_quast/run_QUAST_app/release) to compare the contig distributions of Assembly objects.&#x20;

## Does KBase support co-assembly?&#x20;

Yes, but the results are unstable currently above 200 M reads (Illumina 150bp x 2). Use the [Merge Reads Libraries App](https://kbase.us/applist/apps/kb_ReadsUtilities/KButil_Merge_MultipleReadsLibs_to_OneLibrary/release) to get a combined Reads Library object.&#x20;

## JGI/IMG recommends KBase for genome assembly before submission to them for annotation. Is there a direct way to import to JGI or do we have to save FASTA files and separately submit to JGI/IMG?

These [instructions on transferring data from JGI to KBase](https://docs.kbase.us/data/jgi-transfer) takes you through the process. You can also start from the KBase search and use the JGI tab.

## What is the difference between RAST and Prokka annotations?&#x20;

In addition to having different options in the app, their method for assigning the annotation is different. The determination of better or worse is in the eye of the beholder. The primary advantage of RAST is its linking to our metabolic modeling. The RAST functional roles are considered a controlled vocabulary where we map specific RAST annotations to biochemical reactions in the model, so if you plan to build metabolic models, you should annotate with RAST. Because RAST tends to assign more hypothetical proteins, some people will run Prokka first, and then reannotate with RAST using the Retain old annotation for hypotheticals option.&#x20;

## What is the difference between Annotate Microbial Genome and Annotate Microbial Assembly?

[Annotate Microbial Assembly](https://kbase.us/applist/apps/RAST_SDK/annotate_contigset/release) takes an Assembly object, follows it by gene calling using algorithms from Prodigal and Glimmer, and then functional annotation.&#x20;

[Annotate Microbial Genome](https://kbase.us/applist/apps/RAST_SDK/reannotate_microbial_genome/release) takes a Genome object, does not call genes, and instead preserves the original gene calls. It then re-annotates the genes, overwriting previous annotations.&#x20;

## Is it possible or necessary to manually curate annotations, or are the RAST and Prokka annotations sufficient?&#x20;

Manual curation of annotations is not supported on-system. RAST and Prokka are likely sufficient for many applications, but as you mention, for difficult-to-annotate or highly divergent metabolic genes you may need to use additional tools. In addition to RAST and Prokka, there are on-system tools available for feature annotation using pre-generated hidden Markov models, which could be useful for higher-resolution annotation.&#x20;

## Does RAST annotate archaea and protists?&#x20;

[RAST](https://kbase.us/applist/apps/RAST_SDK/annotate_contigset/release) primarily annotates bacteria and archaea, and may be limited with protists.&#x20;

## Can I annotate plants or fungi?&#x20;

Tools available for plant annotations include [Annotate Plant Transcripts with Metabolic Functions](https://kbase.us/applist/apps/kb_plant_rast/annotate_plant_transcripts/release) and [Annotate Plant Enzymes with OrthoFinder](https://kbase.us/applist/apps/kb_orthofinder/annotate_plant_transcripts/release). These tools use the PlantSEED curated Database.&#x20;

For fungi, it is recommended to use external tools and then import the annotated genomes into KBase.&#x20;

## Are there tools specialized for fungal data?

Currently there is a tool to construct draft metabolic models of fungal species ([Build Fungal Model](https://kbase.us/applist/apps/kb_fungalmodeling/built_fungal_model/release)) that uses highly curated published fungal models as the underlying biochemistry data.


# Comparative Genomics & Phylogenetic Analysis

Running comparative and phylogenetic Workflows in KBase

KBase’s comparative genomics and phylogenetic analysis [tools](https://kbase.us/applist/#Comparative%20Genomics) enable researchers to understand evolutionary relationships between organisms and explore structural and functional variance across genomes. By integrating public datasets with user data and tools to search for sequence homology, explore gene orthology, and construct phylogenetic linkages between organisms, KBase provides an advanced and useful platform for comparative genomics research.

## Narrative Tutorials

* [Build a Gene Tree ](https://narrative.kbase.us/narrative/ws.22290.obj.1)– guides the user through the process of building a gene tree for the "Sulfate Adenylyltransferase" (Sat) gene from Candidatus Desulforudis audaxviator MP104C
* [Phylogenomics of Sulfate Reducing Clostridia - Tutorial - Part 1: Comparative Functional Assignment](https://narrative.kbase.us/narrative/ws.18988.obj.1)– guides the user through the process of whole-genome phylogeny, homology, and domain family functional profiling.
* [Genome Analysis 2: Identifying Features of Interest in Genomes](https://narrative.kbase.us/narrative/83681) – shows how to build a workflow for searching and identifying features within your genomes
* (coming soon) Genome Analysis 3: Domain Annotation and Feature Profiling – guides users through annotating protein families and analyzing features across a set of genomes for comparative analysis&#x20;

## Video Tutorials

{% embed url="<https://youtu.be/XJsWul-lUaU>" %}
Phylogenomics in KBase
{% endembed %}

{% embed url="<https://youtu.be/c6NYWscs_Cw?si=3-A49n1gcphXw1IY>" %}
Microbial Genomics in KBase - Features and Genes of Interest
{% endembed %}


# FAQ: Comparative Genomics

## What does the `tail <1% each taxon` mean in the Kaiju output?

This option allows users to exclude rare lineages which would congest the plot. However, underlying analysis is not affected and the data can be replotted.

## Can I use the GTDB database within Kaiju?

Yes, you can run the our GTDB-tk Classify app on the same data as you run through Kaiju and compare. The GTDB-tk app is in the Microbial Communities section of the Apps Panel in the Narrative interface. One distinction: Kaiju runs on raw sequencing data (reads) and GTDB-tk works on assembled contigs.

## Does the Kaiju analysis give a text output that can be used in R, for example?

Yes, at the bottom of the App cell you can see an option to download kaiju\_classification.zip and kaiju\_summaries.zip which both contain text files.


# Metagenomic & Community Analysis

Running Metagenomic and community analysis in KBase

KBase has a multitude of tools designed to analyze [microbial communities](https://kbase.us/applist/#Microbial%20Communities) and [metagenomic reads](https://kbase.us/applist/#Comparative%20Genomics) data. They provide a way to understand the functional interactions between species in microbial communities. Annotated genomes extracted from metagenomic samples can be used for metabolic modeling, comparative phylogenomics, functional profiling, and more.

## Narrative Tutorials

* [Genome Extraction from Shotgun Metagenome Sequence Data](https://narrative.kbase.us/narrative/33233) – this is a good example for understanding a metagenomic pipeline.
* [Using Metagenomes to Discover Novel Microbial Linages in KBase](https://narrative.kbase.us/narrative/64677) – demonstrates the application of bioinformatic tools for metagenome assembly and genome binning. This Narrative highlights the effective use of employing multiple genome binners and a bin optimization approach.&#x20;
* [Metagenome-Assembled Genome Extraction from a Compost Microbiome Enrichment](https://narrative.kbase.us/narrative/33233) – demonstrates an in-depth example of metagenome-assembled genome extraction and analysis to obtain high-quality genomes.
  * Chivian et al. (2023) *Metagenome-assembled genome extraction and analysis from microbiomes using KBase*. Nature Protocols 18, 208-238. [https://doi.org/10.1038/s41596-022-00747-x](https://rdcu.be/cZJQm)
  * Compare and contrast the enriched environment sample analysis to [a Moab Desert Crust](https://narrative.kbase.us/narrative/62384) environmental sample.&#x20;
* [The PlantSEED Resource in KBase](https://narrative.kbase.us/narrative/39144) – shows the streamlined process of annotating plant genome sequences, automatically constructing metabolic models based on genome annotations, and using models to test annotations.

{% embed url="<https://youtu.be/5K4fZinIOsM>" %}
Microbiome Fractionation by Genome Extraction
{% endembed %}

{% embed url="<https://youtu.be/Poa4cW3hXU8>" %}
Functional and Taxonomic Profiling of Metagenome-Assembled Genomes
{% endembed %}

{% embed url="<https://youtu.be/unFFU3iu92s>" %}
Metagenome-assembled Genome Extraction and Analysis Tutorial Webinar
{% endembed %}


# FAQ: Metagenomics & Community Analysis

## I am very new to shotgun metagenomics based assembly and annotations, and there are many apps are listed in KBase. Does KBase have a pre-developed workflow for shotgun metagenomics, starting from assembly, annotations and metabolic pathways mining?

KBase tries to be as flexible as possible, so there are many options. One App that you could consider is the JGI Metagenome Assembly App. It is a beta app with a complete workflow, optimized by the JGI, that goes from raw reads to an assembly using BFC, BBTools for read QC and metaSPAdes. The assembled sequence can be binned using the available binning tools. And the individual bins annotated using the standard prokaryotic annotation Apps (Prokka, RAST).&#x20;

For annotating the complete metagenome assembly, the Prokka App is being updated to allow this, but it is still in beta. For metabolic modeling of an entire community, there is currently one app, Build Community Metabolic Model that requires a set of genomes or bins as input. It doesn't take the entire metagenome annotation as input since it attempts to model the individual members and the transfer between them. It is possible to make a mixed bag model using the entire metagenome annotation, that can be useful to see entire pathways are present in the metagenome.

## Why extract genome sequences from metagenomes rather than working with unassembled genome sequences?

When you have a contiguous fragment of a genome, there will be 1) full-length genes and their protein products, 2) genomic context of the genes \[to have a better chance of understanding of which genes are being used as part of the same system/pathway, especially if they are polycistronic (operons)], 3) more accurate phylogenetic placement with consensus placement from multiple genes, and possibly even a clade-specific phylogenetic marker.

## What is the best way to assemble 100-150bp sequencing data in order to recover MAGs?

Effective MAG recovery is highly dependent on your sample. If there is a lot of diversity in the sample and/or low read coverage, MAGs are more challenging to recover, regardless of the tools used. This is partly why we recommend evaluating your data prior to assembly, so you can get some idea of what your data look like.

## What causes the contamination in the bins? And what is considered a high quality bin for filtering them out?

Contamination in the bins can occur for a variety of reasons, including but not limited to contig mis-assembly, limited diversity in kmer space, horizontal gene transfer. In general, a genome that is 90% complete and <5% contaminated is high-quality. A rough guide to the quality of MAGs and SAGs can be found [here](https://www.nature.com/articles/nbt.3893).&#x20;

## How would I exclude eukaryotes and viruses from a marine metagenome?

Note that this is hard to do, not just in KBase, but in general. To do this 1) annotate the entire metagenome assembly, 2) identify contigs that fall into the different categories of interest (bacteria/archaea/viruses/eukaryotes), 3) filter out contigs belonging to eukaryotes/viruses, 4) summarize remaining results. A major challenge here will be the unambiguous identification of the different domains of life, which is sometimes tricky (e.g. prophages). Another note: file manipulation outside of KBase would be required to perform this task - as currently there are no KBase Apps to complete this task.


# Transcriptomic Analysis

Running RNA-seq analyses pipelines in KBase

KBase offers a powerful suite of [expression analysis tools](https://kbase.us/applist/#Expression). Starting with short reads, you can use the tool suite to assemble, quantify long transcripts, and identify differentially expressed genes. You can also compare the expression data with the flux when studying metabolic models in KBase and identify pathways where expression and flux agree or conflict.

{% hint style="info" %}
**Prerequisites**

KBase requires a reference genome to guide the analysis of short reads.&#x20;

1. [**Import Genome**](/data/upload-download-guide/genome)
2. [**Import Short Reads**](/data/upload-download-guide/reads)**:** The reads must be a set of single-end, paired-end, or interleaved paired-end reads in FASTA, FASTQ, or SRA format.
3. [**Create a SampleSet**](/data/upload-download-guide/sampleset)**:** Run the [Create RNA-seq Sample Set](https://narrative.kbase.us/#catalog/apps/KBaseRNASeq/describe_rnaseq_experiment/release) App to group together your reads into an RNA-seq sample set with associated experimental metadata to run RNA-seq Apps in batch mode wherever appropriate.
4. [**QC SampleSet**](/apps/analysis/expression#reads-management)**:** Run [FastQC](https://narrative.kbase.us/#appcatalog/app/kb_fastqc/runFastQC/release) to assess the read quality of the reads set from the previous step and if needed, run [Trimmomatic](https://narrative.kbase.us/#appcatalog/app/kb_trimmomatic/run_trimmomatic/release), [Cutadapt](https://narrative.kbase.us/#appcatalog/app/kb_cutadapt/remove_adapters/release), or [PRINSEQ](https://narrative.kbase.us/#appcatalog/app/kb_PRINSEQ/execReadLibraryPRINSEQ/release) to pre-process or filter the reads before starting RNA-seq analysis.
   {% endhint %}

### RNA-seq Pipeline

The RNA-seq pipeline in KBase is modular and consists of three steps. You can pick any of the multiple Apps available for a given step depending on your preference or individual characteristics of the App.

1. **Read Alignment:** [Align reads](/apps/analysis/expression#reads-alignment) to map short reads to the reference genome. The output is a set of BAM alignments and Qualimap report. You can download the alignment output object generated by aligner Apps for further analysis.
2. **Transcriptome Assembly and Quantification:** [Assemble aligned reads](/apps/analysis/expression#reads-assembly) to generate full-length transcripts and quantify transcripts and genes as appropriate. You can view downloadable normalized full expression matrices in FPKM (fragments per kilobase of exon model per million mapped reads) and TPM (transcripts per million).
3. **Differential Gene Expression:** Generate gene- or transcript-level [differential expression](/apps/analysis/expression#differential-expression) based on the quantification. Run [Create Feature Set/Filtered Expression Matrix From Differential Expression](https://narrative.kbase.us/#appcatalog/app/FeatureSetUtils/upload_featureset_from_diff_expr/release) after selecting appropriate q-value and fold change cutoffs as input parameters for the filtering of the differential gene expression.

### Downstream Expression Analysis

1. **Filtering:** You can [create a filtered expression matrix and associated feature set](https://narrative.kbase.us/#catalog/apps/FeatureSetUtils/upload_featureset_from_diff_expr/release) based on fold-change or adjusted p-value. You can also [filter an expression matrix](https://narrative.kbase.us/#catalog/apps/CoExpression/expression_toolkit_filter_expression/release) based on LOR or ANOVA.
2. **Clustering:** Depending on preference, run the [Hierarchical](https://narrative.kbase.us/#catalog/apps/KBaseFeatureValues/expression_toolkit_cluster_hierarchical/release), [K-Means](https://narrative.kbase.us/#catalog/apps/KBaseFeatureValues/expression_toolkit_cluster_k_means/release) or [WGCNA](https://narrative.kbase.us/#catalog/apps/CoExpression/expression_toolkit_cluster_WGCNA/release) clustering App to group features into clusters based on gene expression. You can also visualize the clusters as an interactive heatmap.
3. **Functional Enrichment:** [Assess the functional enrichment](https://narrative.kbase.us/#appcatalog/app/kb_functional_enrichment_1/functional_enrichment_go_term/release) in plant genomes for a set of features using associated GO terms.
4. **Integration into Metabolic Models:** Assimilate the expression data from RNA-seq into the metabolic models to [compare reaction fluxes with gene expression](https://narrative.kbase.us/#appcatalog/app/fba_tools/compare_flux_with_expression) and thus identify pathways where expression and flux agree or conflict.

## **Narrative Tutorials**

* [*E. coli* RNA-seq Analysis Tutorial](https://narrative.kbase.us/narrative/ws.50093.obj.1) – bacterium-based example of an RNAseq workflow using a HISAT2/StringTie/DESeq2 pipeline
* [*Arabidopsis* RNA-seq Analysis Tutorial ](https://narrative.kbase.us/narrative/ws.19391.obj.1)– plant-based example of an RNAseq workflow using a HISAT2/StringTie/DESeq2 pipeline
* [Case Study: Genome-wide Transcriptomics and Plant Primary Metabolism in response to Drought Stress in Sorghum](https://kbase.us/n/101788/79/) – plant-based example of using KBase to integrate full genome RNA-seq analysis and a metabolic model to generate a reaction matrix&#x20;
  * Kumari et al. (2021) *A KBase case study on genome-wide transcriptomics and plant primary metabolism in response to drought stress in Sorghum*. Current Plant Biology 28. <https://doi.org/10.1016/j.cpb.2021.100229>

## **Video Tutorial**&#x20;

{% embed url="<https://youtu.be/mkYA5Ws9UZk>" %}
Functional Genomics&#x20;
{% endembed %}


# FAQ: RNA-seq Analysis

## Can I assemble a transciptome for a species that does not have a reference genome using KBase?&#x20;

Currently KBase only offers RNA-seq using the Tuxedo suite, which usese reference genome-guided transcriptome assembly. We do not have stand-alone tools like Trinity for *de novo* transcriptome assembly.

## Where can I find reference genomes?&#x20;

Public genomes are available through the public data tab. You can directly import genomes from NCBI ReqSeq. You can also import genomes if they are not found in RefSeq. You can use the Upload From Web App if the files are available publicly online, or upload via Globus or the drag and drop upload if they are on your local machine.

## I'm getting an error stating "Failed to generate HISAT2 index files!" or "Warning: Encountered reference sequence with only gaps" during alignment.&#x20;

Check that the reference genome correctly uploaded into KBase. It is possible to export only the features from NCBI which will create a genome object with genes but no corresponding DNA sequences. This prevents the alignment from working correctly.&#x20;

When downloading from NCBI, download the "full" record option to ensure it contains both genes and full-genome DNA sequences.

## Are there some references to what tools (aligner, assembler, etc) we should apply to certain situations (e.g. transciptomics, metatranscriptomics, extreme-conditions communitites)? Or should we apply every combination of tools to every analysis we perform, and see which combination performs better?

&#x20;Certain tools are recommended only for prokaryotes, such as Bowtie2. We recommend HISAT2 for aligner, StringTie for assembler, and DEseq2 for differential gene expression for prokaryotes. We do offer an additional tool, Cuffdiff, for differential gene expression. For eukaryotes, we offer Tophat2 or HISAT2 for assembly and Ballgown, Cuffdiff, and DESeq2 for differential gene expression. Our suggested pipelines are below: For prokaryotes - Bowtie2/HISAT2 ? StringTie ? DESeq2 (DESeq2 is based on gene counts). If you want to compare it with gene abundance, you can run Cuffdiff instead of DESeq2 at the last step. For Eukaryotes - HISAT2 ? StringTie ? DEseq2/Ballgown (you can run with Ballgown only if you want to compare these two different algorithms for differential gene expression). Though Cuffdiff can also work with eukaryotes but it is quite slow and too compute intensive for eukaryotes. So, Cuffdiff is only recommended if you have very few samples.

## When you have replicates, is the analysis done for each of the replicates, or are the replicates pooled to do one analysis?&#x20;

The RNA-seq apps such as aligners (HiSat, TopHat2, Bowtie) and transcriptome assemblers (StringTie, Cufflinks) work on each of the replicates and generate corresponding analysis output objects. However, these apps also wrap the individual results into a set to pass conveniently for the downstream app. However, the Apps such as DESeq2 and ?Create Feature Set/Filtered Expression Matrix From Differential Expression? by definition need to operate on two or more conditions and operate on pools of replicates as opposed to individual replicates.

## Can I compare RNA sequencing data with genome-scale metabolic modes?

&#x20;Yes, the Compare Flux with Expresson allows you to compare the two. You can also use the ?Run Flux Balance Analysis? app to predict fluxes using your expression data. In this approach, you pick a threshold expression level (as a percentile), and the app tries to turn off reactions associated with genes expressed below the threshold and turn ?on? reactions associated with genes expressed above the threshold. We are actively working on this method now, but you are welcome to try it out. It?s fully deployed. Just load you expression data into KBase using the upload tool or by running the RNA-seq pipeline (which will make a matrix for you). Then select this matrix in the advanced parameters of the Run Flux Balance Analysis app, and specify which column of the matrix you want to use to constrain your fluxes.

## Does KBase support metatransciptomic assembly?

&#x20;Cufflinks and Stringtie are available for metatranscriptomics.


# Constructing Metabolic Models

Running metabolic modeling workflows in KBase

KBase has a suite of [Apps](https://kbase.us/applist/#Metabolic%20Modeling) and data that support the reconstruction, prediction, and design of metabolic models in bacteria, fungi, and plants using functionality from [ModelSEED](https://modelseed.org/) and [PlantSEED](https://modelseed.org/genomes/Plants) databases. Genome-scale metabolic models are primarily reconstructed from protein functional annotations. These genome-scale metabolic models can be used to explore an organism’s growth in specific media conditions, determine which biochemical pathways are present, optimize production of an important metabolite, identify high flux pathways, and more.

Metabolic models can be used to evaluate an organism’s metabolic capability by simulating growth under different conditions to answer important biological questions such as:

* What biochemical pathways are present?
* What are the high flux pathways under a certain growth condition?
* Could the organism grow anaerobically?
* Would it grow under certain minimal media conditions?
* Could the organism be optimized to produce a particular drug molecule or industrially important biofuel?

## Workflow

Flowchart of [Apps](https://kbase.us/applist/#Metabolic%20Modeling) used in [Metabolic Modeling](/apps/analysis/metabolic-modeling).

![](/files/-LwQcX23JgDuAhuEVeuY)

{% hint style="info" %}
**Learn More**

The flowchart above shows KBase’s metabolic modeling tools (green) as well as some other analysis tools. Check the [App Catalog](https://kbase.us/applist/#Metabolic%20Modeling) for the latest set of metabolic modeling analysis tools in KBase.

The [video tutorial](https://www.youtube.com/watch?v=AQ2KsrQrq9s\&list=PLh7Q4SqpZYTwdK8ekQnqKinFzbqZuzu8f) below presents an introduction to building metabolic models in KBase. The [“Microbial Metabolic Model Reconstruction and Analysis” Narrative tutorial](/workflows/metabolic-models#narrative-tutorial) lets you see and try out for yourself some examples of KBase’s metabolic modeling functionality in action. Common questions and answers about KBase’s metabolic modeling tools can be found in the [Metabolic Modeling FAQ](/workflows/metabolic-models/faq-metabolic-modeling).
{% endhint %}

## **Narrative Tutorials**

* [MS2 Tutorial - Build and gap-fill genome-scale metabolic models with ModelSEED2](https://narrative.kbase.us/narrative/158450) – Guides through a workflow on how to draft and gapfill metabolic models using improved ModelSEED2 with OMEGGA. [Pre-print.](https://www.biorxiv.org/content/10.1101/2023.10.04.556561v1.article-metrics)&#x20;
  * [Microbial Metabolic Model Reconstruction and Analysis Tutorial](https://narrative.kbase.us/narrative/ws.18302.obj.61) – Guides through a workflow on how to draft a metabolic model and use KBase Apps to gapfill, run flux balance analysis and simulate growth. *Note: Great to reference, but uses an outdated version of ModelSEED.*
* [Microbial Metabolic Modeling in KBase: Drafting & Gapfilling Metabolic Models](https://narrative.kbase.us/narrative/231981) – Demonstrates how to draft and gapfill metabolic models anf run flux balance analysis for prokaryotes.&#x20;
* [Constructing and Analyzing Metabolic Flux Models of Microbial Communities](/workflows/metabolic-models/metabolic-flux-models) – Provides a case study on current research in modeling the growth and behavior of microbial communities.
* Modeling Central Metabolism and Energy Biosynthesis across Microbial Life: [Publication](http://bmcgenomics.biomedcentral.com/articles/10.1186/s12864-016-2887-8); [Narrative](https://narrative.kbase.us/narrative/ws.15253.obj.1)
* Microbial Community Metabolic Modeling: A Community Data-Driven Network Reconstruction: [Publication](http://onlinelibrary.wiley.com/doi/10.1002/jcp.25428/full); [Narrative](https://narrative.kbase.us/narrative/ws.13807.obj.1)
* Case Study: Genome-wide Transcriptomics and Plant Primary Metabolism in response to Drought Stress in Sorghum – plant-based example of using KBase to integrate full genome RNA-seq analysis and a metabolic model to generate a reaction matrix. [Publication](https://doi.org/10.1016/j.cpb.2021.100229); [Narrative](https://kbase.us/n/101788/79/).&#x20;

### Metabolomics

* [Modeling and Integration of Omics Data in KBase](https://narrative.kbase.us/narrative/55494) – Demonstrates model construction, integration of transcriptomics and metabolomics data, and an example of data representation on metabolic maps based on a Yeast example.
* [Predicting Novel Compounds and Reactions using PickAxe](https://narrative.kbase.us/narrative/55494) – Demonstrates cheminformatics expansion of known metabolites, based on compounds within a metabolic model, and using the PickAxe App to map metabolomics data.

## Video Tutorials

{% embed url="<https://youtu.be/J5WtCv6peX8?si=sIX0QWPpLbnBpAHA>" %}
Metabolic Modeling with ModelSEED2
{% endembed %}


# Constructing and Analyzing Metabolic Flux Models of Microbial Communities

KBase has a suite of tools and data that support the reconstruction, prediction, and design of metabolic networks in individual microbes and microbial communities. A recent review article in *Hydrocarbon and Lipid Microbiology Protocols* discusses and compares approaches to modeling the growth and behavior of microbial communities, and shows how to perform the analyses in KBase. The analysis workflows described in the article are available as Narratives (linked below) that you can copy to your own KBase account and rerun, perhaps changing some of the parameters or adding your own datasets.

### **Abstract**

Here we provide a broad overview of current research in modeling the growth and behavior of microbial communities, while focusing primarily on metabolic flux modeling techniques, including the reconstruction of individual species models, reconstruction of mixed-bag models, and reconstruction of multi-species models. We describe how flux balance analysis may be applied with these various model types to explore the interactions of a microbial community with its environment, as well as the interactions of individual species with each other. We demonstrate all discussed model reconstruction and analysis approaches using the Department of Energy’s Systems Biology Knowledgebase (KBase), constructing and importing genome-scale metabolic models of *Bacteroides thetaiotaomicron* and *Faecalibacterium prausnitzii* and subsequently combining them into a community model of the gut microbiome. We also use KBase to explore how these species interact with each other and with the gut environment, exploring the trade-offs in information provided by applying each metabolic flux modeling approach. Overall, we conclude that no single community modeling approach is better than the others, and often there is much to be gained by applying multiple approaches synergistically when exploring the ecology of a microbial community.

### **Narratives**

* Single-Species Model Reconstruction and Analysis: <https://narrative.kbase.us/narrative/ws.10779.obj.1>
* Mixed-Bag Model Reconstruction and Analysis: <https://narrative.kbase.us/narrative/ws.10786.obj.1>
* Mixed-Bag Model Reconstruction and Analysis (VERSION 2): <https://narrative.kbase.us/narrative/notebooks/ws.15122.obj.2>
* Multi-Species Model Reconstruction and Analysis: <https://narrative.kbase.us/narrative/ws.10824.obj.1>
* Species Interaction Analysis: <https://narrative.kbase.us/narrative/ws.10778.obj.1>

**Reference**: Faria JP, Khazaei T, Edirisinghe J, Weisenhorn PB, Seaver S, Conrad N, Harris N, DeJongh M, Henry CS (2016) [Constructing and Analyzing Metabolic Flux Models of Microbial Communities](http://link.springer.com/protocol/10.1007%2F8623_2016_215). *Hydrocarbon and Lipid Microbiology Protocols*, 10.1007/8623\_2016\_215.

Visit our [Publications page](https://www.kbase.us/research/) for more articles that describe research done using KBase.


# FAQ: Metabolic Modeling

KBase has a suite of analysis apps and data that support the reconstruction, prediction, and design of metabolic networks in microbes and plants. These tools can help advance efforts to optimize microbial production of a certain biofuel, find the minimal media conditions under which that fuel is generated, or predict soil amendments that improve the productivity of plant bioenergy feedstocks.

The [“Microbial Metabolic Model Reconstruction and Analysis” Narrative tutorial](/workflows/metabolic-models#narrative-tutorial) lets you see some of this functionality in action.

Below are answers to common questions about KBase’s metabolic modeling tools.

For more information — including descriptions of the underlying algorithms, inputs, and outputs — see the [Metabolic Modeling Apps](/apps/analysis/metabolic-modeling).

## What *is* Flux Balance Analysis?

One of the best papers on the fundamentals of FBA is the aptly titled ["What is Flux Balance Analysis" by Jeffrey Orth, Ines Thiele, and Bernhard Palsson. ](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3108565/)We highly recommend starting there to understand how FBA works in general.

For a detailed description of Model SEED, which KBase uses for modeling, see [High-throughput generation, optimization and analysis of genome-scale metabolic models by Chris Henry, et al.](https://pubmed.ncbi.nlm.nih.gov/20802497/)

## Can Prokka annotations be used for metabolic modeling?&#x20;

You can use Prokka to annotate, but we recommend using RAST for metabolic models because we use the RAST functional roles as a controlled vocabulary to derive metabolic reactions.

## **What does the KBase** [**Gapfill Metabolic Models App**](https://kbase.us/applist/apps/fba_tools/gapfill_metabolic_model/release) **do?**

Draft metabolic models lack a number of essential reactions due to missing or inconsistent annotations. Some of the most common problems arise from having missing transporters (reactions that move metabolites across cell membranes) because they are difficult to annotate well. Consequently, draft models often are unable to generate biomass on media where the organism typically is capable of growing.

The KBase gapfilling process compares the set of reactions in your metabolic model to a database of all known reactions and attempts to find a minimal set of reactions that, when added to your model, will allow it to grow. The set of reactions (known as the gapfilling solution) can then be integrated into your model, creating a new model capable of growth on the media used in the gapfilling process.

The gapfilling app uses a cost function associated with each internal reaction and transporter to find a solution that uses the fewest reactions to fill all gaps, and it does this without extra knowledge about the organism’s biochemistry.

For more information on gapfilling, see the Henry et al. 2010 [ModelSEED paper](http://www.nature.com/nbt/journal/v28/n9/full/nbt.1672.html) in Nature Biotechnology.

## **How does the underlying gapfilling algorithm work (i.e., what kind of programming formulation does it use)?**

At one time, KBase used a mixed-integer linear programming (MILP) formulation in gapfilling but replaced it with simple linear programming (LP), which minimizes the sum of flux through gapfilled reactions. From extensive experience with both formulations, KBase modeling experts have found that LP solutions are just as minimal as MILP solutions but require far less time to compute. In rare cases where an LP gapfilling would not be minimal, the lower wait time to obtain and adjust a new solution is typically worth a slightly larger solution.

The LP insists on using minimal flux, which almost always corresponds with minimal reactions in a stoichiometrically consistent database. In fact, inefficient solutions were more likely to result from the former MILP formulation because KBase sometimes would cut off the solver before an optimal solution was found in order to limit run-times to 24 hours.

In gapfilling, it is important to recognize that not all reactions are created equal. For example, transporters and non-KEGG reactions are penalized, along with reactions that have missing structures or unknown deltaG. To understand all the penalties, see the Henry et al. 2010 [ModelSEED](http://www.nature.com/nbt/journal/v28/n9/full/nbt.1672.html) paper in Nature Biotechnology.

The gapfilling app always has reasons for adding the reactions it does, but those reasons admittedly are not always obvious or biologically relevant. KBase is working on an algorithm that would provide detailed explanations of why each gapfilled reaction is added; we hope to make this feature available soon.

## **What is the solver used in your Gapfill optimization?**

The Gapfill optimization uses the [SCIP solver](http://scip.zib.de/).&#x20;

We actually use two solvers in KBase. We use GLPK for most pure-linear optimizations. We use SCIP for larger more complex problems, particularly when integer variables are involved. Gapfilling is an example of a problem where we use SCIP.

## **What is “complete” media?**

\
The “Complete” media is an abstraction of what’s available in our biochemistry database. Every compound that can be transported into the extracellular compartment–or, in other words, for which a transport reaction is available–is used in the complete media. This list is built in real-time, meaning that whenever you run FBA with complete media, the available transporters are parsed from the media database–and is therefore not stored permanently in any media object in the workspace.

To find out which transport reactions are available when running FBA using Complete media, use the [KBase ModelSEED Biochemistry Database ](https://github.com/ModelSEED/ModelSEEDDatabase/tree/v1.0/Biochemistry)for a reference list of reactions and compounds or [Biochemistry Search](https://narrative.kbase.us/#biochem-search) within KBase.

To see the specific compounds that got consumed/excreted when a model was gapfilled in complete media, you can run FBA on complete media. In the ‘FBA output table’ there is a tab called “Exchange Fluxes” which gives a list of excrete/uptake compounds. You can also view the gapfilled reactions in the ‘model output table’ by choosing the “Gapfilling” tab.

## **How do I choose a media condition for gapfilling, and what is the default media used?**

“Gapfilling” refers to choosing the minimal set of reactions to add to a draft metabolic model that will enable it to produce biomass in a specified media. KBase’s gapfilling app lets you set the media condition (i.e., the metabolites available in the environment in which you want to analyze the growth of your organism). For more information, see the [Gapfilling App details page](https://kbase.us/applist/apps/fba_tools/gapfill_metabolic_model/release) and the previous question.

If you leave the Media field blank, “complete” media will be used by default. Please see the next question for more details about complete media. When used on complete media, the gapfilling app is not limited by the transport reactions already available in the model and therefore invariably will add more transport reactions to the model.

In addition to the default complete media, KBase provides more than 500 media conditions you can use for gapfilling. The [Narrative Interface User Guide](/getting-started/narrative) describes how to add these media and other data to your Narrative. You can also [upload your own custom media condition](/data/upload-download-guide/media).

Choosing minimal media for the initial gapfilling is often a good idea because it ensures that the gapfilling algorithm will add the maximal set of reactions to the genome-inferred model. This will allow the model to biosynthesize many common substrates necessary for growth—substrates that otherwise would be present in the media. The gapfilling algorithm looks for incomplete metabolic pathways and tries to make an informed guess as to which protein-encoding genes may be responsible for the missing step in the pathway. This process works best when informed by the a priori knowledge that a particular organism can grow on a defined media type (e.g., a model of *Helicobacter pylori*, an endosymbiont, might require a media containing substrates it can’t biosynthesize in vivo.) Thus, specifying a growth media is beneficial.

Note that gapfilling runs can be stacked, meaning that multiple gapfilling solutions are incorporated into the same model, one after the other. In other words, if you gapfill on Complete media first (i.e., don’t specify a media, so it defaults to Complete), gapfilling will add all reactions needed to get the model to grow assuming it can transport all compounds for which a transporter is available in the KBase biochemistry database. If you then gapfill the same (already gapfilled once) model on minimal media, the model will only add any additional reactions needed to grow on the minimal media.

If you want to gapfill on the two media independently, be careful not to overwrite your original model during your gapfilling simulations and then use the same original model in both gapfilling runs.

The [tutorial for the Microbial Metabolic Modeling and Analysis of *Shewanella oneidensis*](https://narrative.kbase.us/narrative/18302) explores the differences between two gapfilled solutions: one generated on complete media, and one generated on minimal media.

## **How can I see the compounds comprising a media used in gapfilling?**

To see the list of media compounds for a particular model, open a viewer for the model by either clicking on its name in the Data Panel or dragging it from the panel and dropping it into the main Narrative window. Select the Compounds tab and filter by “e0” (extracellular compartment). This gives you a complete list of compounds transportable by your model.

## **Which reactions were added in the gapfilling process and why were they added?**

After performing gapfilling on a metabolic model, you can sort the “Reactions” tab of the output table by the “Gapfilling” column to see which reactions were added. In order to compare the differences between the draft model and the gapfilled model, look at the directionality of the reaction in the “Equation” column of the “Reactions” tab of the output table. If the directionality is reversible in the gapfilled model, i.e. “<=>”, this means that the reaction was already in the draft model, albeit previously irreversible, and gap-filling made it reversible. However, if the reaction is irreversible e.g., “=>” or “<=”, then this means it is a new reaction that wasn’t previously present in the draft model, and has been added to the model by the gapfilling algorithm.

The primary function of gapfilling is to ensure that the model in question can produce the biomass of the organism. As gapfilling itself is a heuristic, the result of the algorithm is essentially a prediction and requires manual curation. If the addition or reversibility of a reaction is not desired, you can use the “Custom flux bounds” field to force the reaction to be zero, and re-run gap-filling to find another possible solution. For the gapfilling process, it does not necessarily matter whether a reaction was previously set to be irreversible as long as it produces biomass. However, even if it was determined that a reaction was irreversible *a priori*, the thermodynamics algorithm used to determine this is also a heuristic.

## What conventions are used for metabolic IDs, reaction IDs, etc?&#x20;

If you generate a metabolic model based on the [Build Metabolic Model App](https://kbase.us/applist/apps/fba_tools/build_metabolic_model/release), the biochemistry represented is based on modelSEED. However, if you generated a model based on a published model and propagate the biochemistry into a new model based on proteome comparison data you could have a mix of ModelSEED and the native reaction and compound IDs from the published model. We recommend using the ModelSEED ontology; you can convert a mixed model using the [Integrate Imported Model into KBase Namespace App](https://narrative.kbase.us/#catalog/apps/SBMLTools/integrate_imported_model/beta) (*in beta*).

## **When I import a metabolic model, how do I associate it with a genome?**

Load the genome you wish to associate with a metabolic model into your Narrative. Next, when importing your metabolic model, click on the “show advanced options” link shown towards the bottom of the import metabolic model menu. When you click on this link, a new input will appear that will permit you to select the genome for your model. Refer to the [FBA Model section of the Data Upload and Download Guide](/data/upload-download-guide) to learn more about correctly importing and associating your model and genome.

## **When viewing a model, what do the compartment IDs mean?**

\
Compartments are the subcellular localization of compounds, enzymes, and reactions. The table below lists compartment abbreviations for microbes and plants.

Reactions and compounds belonging to each compartment are identified using compartment notation in square brackets after the five-digit reaction or compound ID (e.g., rxn00001\[c0], cpd00001\[c0]).

The integer associated with the compartment (e.g., the 0 in c0) represents the index number of the model. For a single-species model, this number will always be zero, but if individual models are merged into a community model, each sub-model will then be assigned a distinct index.

| **Compartment**           | **Microbes** | **Plants** |
| ------------------------- | ------------ | ---------- |
| Cytosol (c)               | ✓            | ✓          |
| Plastid (d)               |              | ✓          |
| Extracellular (e)         | ✓            | ✓          |
| Golgi apparatus (g)       |              | ✓          |
| Mitochondria (m)          |              | ✓          |
| Nucleus (n )              |              | ✓          |
| Periplasm (p)             | ✓            |            |
| Endoplasmic reticulum (r) |              | ✓          |
| Vacuole (v)               |              | ✓          |
| Cell wall (w)             |              | ✓          |
| Peroxisome (x)            |              | ✓          |

## **After merging models, why do I sometimes end up with one or two more compartments?**

When you combine models into a community model, you have two options:

1. Make a compartmentalized multi-species model where each species is contained in a separate compartment leading to the replication of compartments that you describe
2. Make a “mixed bag” model where the component models are merged non-redundantly into a single model with no additional compartments.

The first model type is better for predicting potential interactions between species. The second model type is a simpler model and at times, easier to work with. These papers and the narratives associated with these papers explain this in greater detail:

* Henry CS, Bernstein H, Weisenhorn P, Taylor RC, Lee JY, Zucker J, Song HS. Microbial Community Metabolic Modeling: A Community Data-Driven Network Reconstruction. Journal of Cellular Physiology (2016) 10.1002/jcp.25428.
* Faria JP, Khazaei T, Edirisinghe J, Weisenhorn PB, Seaver S, Conrad N, Harris N, DeJongh M, Henry CS. Constructing and Analyzing Metabolic Flux Models of Microbial Communities. Hydrocarbon and Lipid Microbiology Protocols (2016)

## **What should I know about performing metabolic modeling on plants?**

The Plant Metabolic Modeling in KBase is based on the metabolic subsystems described in the [PlantSEED](http://bioseed.mcs.anl.gov/~seaver/FIG/seedviewer.cgi?page=PlantSEED). Because PlantSEED was focused on the primary metabolism of higher plants, performing metabolic modeling on lower plants such as bryophytes and algae is not recommended. In principal, the majority of the metabolic networks for plants will be similar. What will be different in the plant primary metabolic network between organisms the gene – reaction associations. From metabolic modeling, you can derive a prediction of the enzymes that would catalyze each of the reactions in the primary metabolic network.

## **How can I find and download all the biochemical compounds and reactions in KBase, as well as linked information such as Enzyme Commission (E.C.) Numbers?**

The biochemistry compounds and reactions can found in the [KBase ModelSEED Biochemistry Database ](https://github.com/ModelSEED/ModelSEEDDatabase/tree/v1.0/Biochemistry)for a reference list of reactions and compounds or using [Biochemistry Search](https://narrative.kbase.us/#biochem-search). This spreadsheet includes the EC numbers. You can also consult the [Enzyme database](http://enzyme.expasy.org/) to see a description of each type of characterized enzyme that has an EC number.

## **Does media in KBase include sources other than carbon (e.g., nitrogen, phosphorus, sulfur)?**

Every KBase media has all the compounds necessary for minimal growth. In most cases the significant difference among media is the carbon source, which generally dictates the media’s name (e.g., “Carbon-D-Glucose”).

## **How can I specify my own media?**

You can generate and import your own media file by specifying the biochemical compounds found in a particular growth environment. Go to  [KBase ModelSEED Biochemistry Database ](https://github.com/ModelSEED/ModelSEEDDatabase/tree/v1.0/Biochemistry)to see a list of all the biochemical compounds in KBase and find the IDs for those you wish to include in your custom media.

Media can be uploaded as a TSV (tab-separated values) or Excel file with four columns:

* Compound identifier (e.g., ModelSEED ID, KEGG ID, PubChem ID, or compound name such as “glucose”)
* Concentration (concentration of compound in mol/L)
* Min flux (minimum allowed uptake/excretion of compound)
* Max flux (maximum allowed uptake/excretion of compound)

Below is an example of a media file in TSV format. In this case, the compound IDs in the first column are ModelSEED IDs.

Note that when creating a media Excel file, the name of the worksheet that contains your media conditions must be named “MediaCompounds” or it will not upload. To rename a worksheet, find the worksheet tabs at the bottom left of the Excel file. Double-click on the default name of the spreadsheet containing your custom media and name it “MediaCompounds."

The values in the “minflux” and “maxflux” columns are absolute units that represent the range of possible fluxes for reactions that transport the media compounds into and out of the cell. A negative flux value corresponds to excretion (transport out of the cell), and a positive flux value corresponds to uptake (transport into the cell).

The combination of maximum and minimum flux will dictate whether uptake and/or excretion of a reagent is allowed. Assigning a positive number to “minflux” for a particular compound essentially forces uptake of that compound. Conversely, assigning a negative number to “maxflux” will force excretion of the compound.

If you specify a range between “minflux” and “maxflux,” then you will need to look at the list of “Exchange Fluxes” in the output after running flux balance analysis to observe what actually happened during the growth simulation.

Please see the [Media section of the Data Guide](/data/upload-download-guide/media) for more information on how to configure your media file.

## **How are units of flux measured within the media?**

Minimum and maximum flux are measured in mmol per gram cell dry weight per hour.

## What is the unit of measurement for the estimated biomass from a metabolic model?

The biomass components are normalized so that a result of 1 (objective value/growth/flux through biomass reaction) is estimated to be 1 gram of dry weight of biomass.

## **How is media data structured in KBase?**

As with all KBase data types, the technical “spec” files of media data can be viewed at[ narrative.kbase.us/#spec/type/KBaseBiochem.Media](https://narrative.kbase.us/#spec/type/KBaseBiochem.Media).

From here, you can find details on the structure of media data in KBase by navigating the following path: “KBaseBiochem” → “Types”→ “Media” → “Spec file.” Scroll to the bottom of the media spec file to see the overall data structure, which should look like this:

```
typedef structure {
  media_id id;
  string name;
  source_id source_id;
  string source;
  string protocol_link;
  bool isDefined;
  bool isMinimal;
  bool isAerobic;
  string type;
  string pH_data;
  float temperature;
  string atmosphere;
  string atmosphere_addition;
  list<MediaReagent> reagents;
  list<MediaCompound> mediacompounds;
} Media;
```

## **What is the function of the .gffba file? What is the difference between the .gffba output and Flux Balance Analysis output?**

The .gffba object is generated automatically by gap filling and it holds the fluxes corresponding to the solution selected by gap filling to make your model grow in the specified environment. It is useful as a means of exploring why various reactions were added, as well as diving into the gap filling result in detail. We are working to improve the functionality of this data object.

The .gffba output and Flux Balance Analysis output are the same. It is important to understand that a gap filing simulation is a form of FBA, and as such, it generates a flux solution, which is exposing to you via the .gffba object.

## **How can I determine if my metabolic model will grow in a certain media?**

Once you have built a metabolic model, you can perform flux balance analysis ( FBA) to calculate the flow of metabolites through your model. FBA results can be used to predict the growth rate of an organism under certain conditions or the production rates for particular metabolites of interest.

To determine whether your organism grew in the media that you specified for FBA, find the Overview tab in the FBA results table. Among the summary information in this tab is the **objective value**. This value is is significant because it represents the maximum achievable flux through the biomass reaction of the metabolic model. An objective value of 0 (or something very close to 0) means that the model did not grow on the specified media.

For more information on the [Run Flux Balance Analysis](https://kbase.us/applist/apps/fba_tools/run_flux_balance_analysis/release) App.&#x20;

## **After running flux balance analysis, what does the sign of a reaction flux mean?**

A positive flux value indicates that a reaction is proceeding from left to right as written in the equation, while a negative value indicates the reaction is running from right to left.

For example, the following table shows a portion of the “Reaction fluxes” tab from the output of the [Run Flux Balance Analysis](https://kbase.us/applist/apps/fba_tools/run_flux_balance_analysis/release) App. Here, the positive flux in the “Flux” column indicates that rxn00543 is proceeding from left to right—that is, ethanol is being oxidized to acetaldehyde.

| **Reaction**  | **Name**                             | **Flux** | **Min flux (Lower bound)** | **Max flux (Upper bound)** | **Class** | **Equation**                                                                  |
| ------------- | ------------------------------------ | -------- | -------------------------- | -------------------------- | --------- | ----------------------------------------------------------------------------- |
| rxn00543\[c0] | <p>Ethanol-NAD<br>oxidoreductase</p> | 100      | <p>-100<br>(-1000) </p>    | <p>100<br>(1000) </p>      | Variable  | <p>NAD\[c0] + Ethanol\[c0] <=><br>Acetaldehyde\[c0] + NADH\[c0] + H+\[c0]</p> |

The “Min flux,” “Max flux,” and “Class” columns show that other valid FBA solutions for the model have flux values for rxn00543 that range from -100 to 100. This reaction therefore is classified as “Variable” because it can proceed in either direction.

## **What do the various reaction classifiers mean?**

Reaction classifiers are assigned when using Flux Variability Analysis, FVA. In FVA, the global objective (biomass) is fixed at its optimal value, then each reaction, iteratively, is optimized independently to find both the maximal and minimal value that is possible given that the global objective must still be reached.

* Variable – the reaction has positive maximal and negative minimal values, meaning that it can go in either direction.
* Positive variable – the reaction has a positive maximal, and a zero minimal, meaning that it can either be zero, or it can go from left to right.
* Negative variable –  the reaction has a zero maximal, and a negative minimal, meaning it can either be zero, or it can go from right to left.
* IA (InActive) – the reaction is blocked and cannot have a non-zero value.

## **Can I customize the flux boundaries for reactions in my model?**

To constrain the flux of a reaction, you can export your model, edit it, and re-import it.

Alternatively, in the case of a transport reaction, you can change the uptake of a media compound by editing the media object itself. For example, setting the maximum flux of any media compound to zero (or a small number) will prevent its uptake. This is contingent, of course, on the media itself being defined, so if you performed gapfilling on the default complete media, you will have to re-gapfill on an actual defined media object of your choice.

Also, in the [Run Flux Balance Analysis](https://kbase.us/applist/apps/fba_tools/run_flux_balance_analysis/release) App, you can change the “bounds” on a specific reaction to whatever you want using the “Custom flux bounds” field under advanced options.

Note that the help text next to the custom flux field provides an example of the proper format for setting bounds:<br>

* Specify a number for the lower bound (e.g. “0”)
* Insert a semicolon (no spaces!)
* Type in the reaction ID or compound ID
* Insert another semicolon
* Specify a number for the upper bound (e.g. “5”)

So to set “rxn00006” at a flux of “5,” you would set this custom bound: 5;rxn00006;5.

## **How can I edit my media or metabolic model (add or delete reactions, compounds, or biomass)?**

After you have drafted your metabolic model with the Build Metabolic Model app, you can use the [Gapfill Metabolic Model App](https://kbase.us/applist/apps/fba_tools/gapfill_metabolic_model/release) to automatically fill in the gaps of necessary reactions for completing the pathway for a set of specified reactions (this can be either the biomass producing reaction or other specified pathways). You can alternatively use our [Edit Media](https://kbase.us/applist/apps/fba_tools/edit_media/release) or [Edit Metabolic Model](https://kbase.us/applist/apps/fba_tools/edit_metabolic_model/release) Apps. This works best if the reactions you're adding at ModelSEED reactions (ID looks like rxn#####). In those cases, you can manually list several reactions to add by ID and associate genes and specify directionality. You can add completely custom reactions as well, but this is harder.

## **Why did gapfilling fail for a model containing a custom media or custom biomass reaction? What can I do about it?**

Draft metabolic models tend to have gaps by default. If you simulate FBA on a given media, you would see the objective value to be zero; means the model is not able to grow on the simulated media. In the model output viewer gapfill reactions can be seen and easy to recognize as there are no features mapped to the gapfilled reactions. There is a tool in KBase to try and find gene candidates for gapfilled reactions. If a candidate is found, this can give you more confidence that the gapfilled reaction is real. It's in beta and it's called [Find Candidate Genes for a Reaction](https://narrative.kbase.us/#catalog/apps/kb_reaction_gene_finder/find_genes_from_similar_reactions/beta). Just plug in the list of gapfilled reactions and see the candidates that come out.

If the [Gapfill Metabolic Model App](https://kbase.us/applist/apps/fba_tools/gapfill_metabolic_model/release) fails with custom media, first check whether the compound that you are gapfilling for is already available in the reconstructed model. (See instructions for viewing model compounds above.)

If the compound is not in the model, it may be rare or uncommon. Check [KBase ModelSEED Biochemistry Database ](https://github.com/ModelSEED/ModelSEEDDatabase/tree/v1.0/Biochemistry)or [Biochemistry Search](https://narrative.kbase.us/#biochem-search) to ensure that:

* The compound is actually available in KBase’s biochemistry collection.
* There are reactions that involve the compound in question.
* There are exchange reactions (transporters) that would enable the compound to be consumed from the media.

To help mitigate this issue, the “Gapfill Metabolic Model” app includes a field for “Source Gapfill Model” under the advanced options. You can specify a source model to supply reactions missing from the KBase database that might prevent the gapfilling from working successfully. If your biomass reaction is from an existing published model, you should import this model and use it as a source database for your gapfilling (because that model should contain all reactions needed to grow on its own biomass reaction). The source model simply supplements the KBase database, which still will be used in gapfilling.Successfully running the gapfilling app with a custom biomass reaction depends on whether the custom biomass includes compounds that can be produced given all the balanced, approved reactions in the KBase database. For now, non-standard compartments for biomass components may also be a problem.

KBase uses a  [KBase ModelSEED Biochemistry Database ](https://github.com/ModelSEED/ModelSEEDDatabase/tree/v1.0/Biochemistry)that includes all ModelSEED reactions, and some additional reactions. However, the gapfilling is restricted to a subset of the reactions in the database. First and foremost, KBase filters out imbalanced reactions from gapfilling because they cause mass leakage and a bad gapfilling solution. We also filter out several reactions that cause problems in gapfilling due to strange formulations (e.g., reactions containing lumped or undefined compounds).

## **What if my source and translated metabolic models are the same after running the “Propagate Genome-scale Model to Close Genome” App (in beta)?**

The [Propagate Model to New Genome App](https://narrative.kbase.us/#catalog/apps/fba_tools/propagate_model_to_new_genome/beta) (in beta) builds a new metabolic model by translating the gene associations in an existing model to a new genome. This translation is conducted based on gene correspondence tables contained in a proteome comparison object generated in Step 1 of the app. A gapfilling is performed to permit the translated model to grow on a specific media. More information on the model translation step can be found on the [details page](https://narrative.kbase.us/#catalog/apps/fba_tools/propagate_model_to_new_genome/beta).

If the translated model resulting from this app is the same as your source model, run through the following checks:

* **Make sure the reactions in the source model have genes associated with them.** If they don’t, then the source and translated models will be the same because reactions with no genes are automatically retained in the model propagation process. For this reason, it is very important to associate a genome with your model when you import it. It is equally important that the gene IDs in this associated genome match the gene IDs mapped to the reactions in your imported model **exactly**. Otherwise, the genes in the genome will not be properly mapped to the reactions in the model. This proper mapping is absolutely essential for apps such “Propagate Genome-scale Model to Close Genome” to work.
* **Look at the proteome comparison object.** How different are the genomes that are being compared? How many hits are there? If the genomes are very close, the source and translated model could end up being the same. Such a result is unlikely but not impossible.
* **Check the translated model before gapfilling.** Did you lose some reactions only to have them added back in the gapfilling process?

## **What do the “Prediction class” values mean in the output table generated by the “Simulate Growth on Phenotype Data” App?**

The [Simulate Growth on Phenotype Data](https://kbase.us/applist/apps/fba_tools/simulate_growth_on_phenotype_data/release) App performs multiple flux balance analyses (FBA) using each media condition in a phenotype dataset and compares the results, which are listed as “Observed normalized growth” (experiments done in a laboratory) vs. “Simulated growth” (simulation of FBA on the model). You can see these results in the Phenotypes tab of the output cell.&#x20;

In the “Prediction class” column, simulated growth results against observed growth are represented in four classes:

**CP** — Correct positive (model was predicted to grow, and did)\
**CN** — Correct negative (model was predicted not to grow, and did not)\
**FP** — False positive (model was predicted to grow, but it did not)\
**FN** — False negative (model was predicted not to grow, but it did)

## What do the different colors mean in the visualization of metabolic pathways?&#x20;

Blue means that the reaction is present, white means the reaction is absent. Red means negative flux, green means positive flux. Shade represents intensity of flux. The arrows on the map do not always agree with the direction of reactions as defined in our biochemistry database, so these colors can appear confusing at times. For this reason, we are actually currently working on replacing our KEGG-based metabolic map viz with a new system using dynamically drawn maps in Escher.

## **Why did I get false negatives after simulating growth, and what can I do about it?**

False negative (FN) results from the “[Simulate Growth on Phenotype Data](https://narrative.kbase.us/#narrativestore/method/simulate_growth_on_a_phenotype_set)” app mean that the model is not able to grow on laboratory-tested media conditions. The most likely explanation for a FN prediction is that the model is missing reactions required for viable growth in the specified media conditions (e.g., no asparagine metabolism pathway in a media where asparagine is the only carbon source). This error can be corrected by gapfilling the model on that media to add a minimal set of reactions to the model such that growth is permitted.

## Why do I get gapfilling that forces unrealistic growth?

It is important to understand that just because gap-filling “works” doesn't mean its results are always accurate when compared to experimental data. The gap-filling algorithm is an optimization problem; given the objective of growing in a particular media/carbon source, the algorithm will search the database for a solution that produces all compounds in the biomass reaction. If it can find a thermodynamically feasible solution in the database, it will add the reactions, even if these solutions can be "unrealistic.” If you look at the gap filling reactions, they don't have genes associated with them, so reactions that don’t reflect enzymes in your organism can be added just because the algorithm found that to be the most optimal solution given the media, model, and biomass reaction provided. Always be inquisitive of gap filling results. Inspect the gap filling results and try to understand them. In this case, you know that growth is impossible, but when growth is possible is still important to understand the gap filling results. A common analogy is that gap filling can be like a zip code instead of an exact address. This means that gap filling can provide you clues of something that might be wrong in a specific pathway but not necessarily add the exact reactions you expect; it may add some alternative reactions that the algorithm ( that, unlike you, does not know biology) deemed to be more efficient as a solution to the optimization problem. It is important to understand why gap filling can happen. Some likely scenarios that lead to gap filling:\
\
1\) The source genome(s) are missing gene function annotations due to poor genome quality or lack of ability from the genome annotation tool to annotate certain functions properly. When this happens, gap filling might add the reactions to fill in missing annotations. Usually, I look at KEGG, Metacyc, UNIPROT, etc., as part of my gap filling analysis, trying to understand if my organisms of interest have the enzyme corresponding to the reactions added by gap filling. As a rule of thumb, if an entire pathway is being gap filled, more likely than not that the gap filling found an optimal solution that is “unrealistic.” But if only one reaction in a pathway is gap filled, that is probably more accurately a reaction missing that should be there. Another example is missing transporters. Transporters are notoriously difficult to annotate, and some times you will see gap filling just adding a missing transporter as the solution because we could not find an annotation for it in the genome.

2\) There is an issue with the model reconstruction template that we provide for model reconstruction. In the backend of our tools, we use a template comprised of associations between gene functional roles annotated with RAST and reaction in the ModelSEEDDatabase. These associations can be wrong or incomplete. We do our best to curate and update our templates, but it takes analyzing hundreds of models to assess if the template is as comprehensive as it needs to be for a multitude of different organisms. Reports like yours with possible issues in this ticket are helpful for me to see if a fix in the template is necessary. For example, looking at your table above, I know that 2 or 3 reactions are related to an issue in our template with thiamin biosynthesis. We know this issue and have a fix in place that will be deployed in a future release (no date for release yet). But the fact that gap filling adds these reactions doesn't make your model “wrong” or “incorrect”. Thiamin is in the model biomass reaction, meaning it needs to be produced for the model to grow. By default, the gap filling goal is to produce all compounds in the biomass reaction, so if reactions are missing to accomplish this for any biomass compounds, gap filling will add them to the model to ensure the production of all biomass-reaction compounds. In this case, for example, if you add Thiamin to the media, those reactions should stop being gap filling, since the model doesn't need to add the missing reactions for biosynthesis; it can just get it from the media directly.

3\)The gap filling adds reactions to produce biomass compounds irrelevant to your specific organism. We have generic biomass reactions, one for gram-negative, gram-positive (bacteria), and one for Archaea. The gram-negative and gram-positive ones only differ in the cell wall compounds necessary for the gram positives species. Its one of the most complex problems to solve, and the current limitations of automated model reconstruction to get as specific as possible biomass reactions for each organism. These are experiments that take months in a lab to perform. Even in published manually curated metabolic models, they usually skip this step and choose to adopt and modify the biomass from E.coli (or another close organism to their organisms where data is available). Meaning the gap filling can unrealistically be adding reactions to produce a biomass compound your organism doesn't need, because its in the biomass reaction, and the objective of the optimization problem is to produce all biomass compounds while achieving the highest growth rate possible. One solution here can be to edit the model to remove non-relevant biomass reaction compounds from your organism.

It is important to understand that we always expect some gap filling due to the lack of specificity of our templates, but that may not impact your modeling analysis. For example, as I mentioned above, thiamin biosynthesis has an issue in our template, and gap filling adds reactions to fix it. But if the goal of your study is to study some pathway or set of pathways represented properly in the model, the thiamin gap filling won't be an issue; it's just necessary to fulfill the biomass reaction growth requirements. If the model produces energy as you expect and the flux distribution for the problem you are trying to solve looks accurate, some level of gap filling in the model is not a problem.

&#x20;

With all that said, it can be hard to assess all of what is described above; this requires knowledge in building and analyzing metabolic models. Our goal is to provide more information that can help users determine what is going on with gap filling and the energy biosynthesis capabilities of models.

## Is it possible to use KBase to integrate RNA sequencing data into existing genome-scale models?&#x20;

Yes, there is an app called [Compare Flux with Expression](https://kbase.us/applist/apps/fba_tools/compare_flux_with_expression/release) that allows you to compare your GSM with RNA-seq data to identify metabolic pathways where expression and flux data agree or conflict. You can also use the [Run Flux Balance Analysis App](https://kbase.us/applist/apps/fba_tools/run_flux_balance_analysis/release) to predict fluxes using your expression data. In this approach, you pick a threshold expression level (as a percentile), and the app tries to turn off reactions associated with genes expressed below the threshold and turn on reactions associated with genes expressed above the threshold. Just load your expression data into KBase using the upload tool or by running the RNA-seq pipeline (which will make a matrix for you). Then select this matrix in the advanced parameters of the [Run Flux Balance Analysis App](https://kbase.us/applist/apps/fba_tools/run_flux_balance_analysis/release), and specify which column of the matrix you want to use to constrain your fluxes.

## What exactly is the criteria for gapfilling/ How do we know which reaction should be maximized, added, etc? Should we consider the reactions which are blocked and do gapfilling for those?

Gap filling is an optimization problem; by default, we gap fill for growth, meaning that the goal of the algorithm is to ensure that given the media provided by the user, all compounds in the biomass reaction (necessary for growth) can be produced by the model. The algorithm searches our biochemistry database and finds reactions that fill gaps that enable the production of the biomass precursors that the draft model without gap-filling can’t produce.

The algorithm tries to minimize the number of reactions that are added to the model to achieve growth since, at the moment, we don’t have a way to map the reactions being added by gap filling to genes with missing or wrong annotations in your genome. This is why gap-filling reactions don’t have gene-protein-reaction associations. If you look at the algorithm in detail in the paper mentioned above, you will see the penalty system we use for adding reactions to a model with gap-filling; for example, thermodynamic feasibility weighs heavily in deciding if a reaction is added to the model or not. Another factor can be the penalty scoring of adding a transporter vs. adding a full biosynthesis pathway if a compound is present in the media.

Reactions can be gap-filled for many reasons:

1. The genome used to build the model has missing or incorrect annotations for a given gene(s), meaning the reaction associated with the gene won’t be placed in the model.
2. Our template is missing or has incorrect mappings between gene functional roles and reactions. While we try our best to have accurate mappings in our templates, knowledge for some pathways keep evolving, EC number get re-assigned, and more specific enzyme complex roles are mapped to individual genes, etc. If you find that a reaction missing in your model should have been included, given that the gene is properly annotated in the source genome, please let us know so we can fix it in our templates.
3. There is truly unknown knowledge of how some organisms perform certain processes.
4. As mentioned above, gap filling will try to make sure all biomass precursors needed for growth are generated by the model. This biomass reaction is mostly generic/universal with some differences in cell wall biomass precursors between gram-negative and gram-positive organisms. Your organism may not have the same biomass requirements as the universal biomass formulation we use, meaning the gap-filling algorithm may gap-fill reactions to produce biomass precursors that are irrelevant to your organism.

If in the gap-filling problem, you decide to maximize the production of a different compound, for example, maximize the production of ethanol by selecting the ethanol transporter instead of the biomass reaction, the gap-filling will act in the same way. Still, this time around will only care to add reactions that fulfill this objective and address gaps in the ethanol production pathway. And important to note that gap filling may add reactions even if your organism cannot produce ethanol. The algorithm does not know the specifics of your organism, but the gap-filling results can help elucidate if this is something or organism can do. It is important to analyze the gap-filling results. Continuing with the previous example. If to produce ethanol in your organism, gap-filling fill only one gap in the pathway, that probably means there was a missing or wrong annotation in your genome, and gap-filling correctly filled the necessary gap. If, on the other hand, gap-filling fills in most of the reactions present in that pathway, it might mean that this is a metabolic capability that your organism doesn't have. Or your source genome has bad quality annotations that miss our reactions mapping completely for this pathway.

TLDR: it is very important to analyze the gap-filling results critically to understand if some reactions are truly missing or are being added to satisfy a biomass precursor that is not relevant to your organism. I also recommend you do further literature research on the subject since multiple gap-filling algorithms exist that employ different methodologies.

## Is there a tutorial, public narrative, publication, or documentation of using a high-quality, curated metabolic model to develop a new metabolic model for a genome of a different strain of the same species?

We have an app Propagate Model to New Genome which can be used to do that. The [app catalog page](https://narrative.kbase.us/#catalog/apps/fba_tools/propagate_model_to_new_genome/release) has good documentation about using the app.

There’s a public Narrative by Chris Henry that you can view as an example here: [![](https://narrative.kbase.us/favicon.ico)KBase](https://narrative.kbase.us/narrative/13806). For context, it is part of this publication: <https://onlinelibrary.wiley.com/doi/full/10.1002/jcp.25428>.

In addition, check out this protocol paper: <https://link.springer.com/protocol/10.1007/978-1-0716-1585-0_13><br>

## I have built a community model on KBase. How can I interpret results from that?

[This narrative](https://narrative.kbase.us/narrative/13838) should provide some starting points of analysis and the interpretation of the data. Primarily researchers are interested in seeing cross feeding events/exchanges across organisms and explain the sustainability of the community against its native environment. The narrative provides an example of a code-cell in identifying these cross feeding events.

## How can I identify as to which organism is showing most benefit and which organism is showing least benefit?

This is a great question, and depends on a lot of factors such as; specific metabolic capabilities of members in the community, their relative abundance, media/environment that they are growing, energy biosynthesis strategies etc. This could be its own little research project to address these important scientific questions.

## How can I determine the biomass coming out from each microorganism?

With the current tools, you can modify the biomass of each member of the community thus fix the biomass based on the relative abundance.

## How can I figure out the metabolites coming from different microorganisms that are a part of the community?

If you know the media that community grows on, any exchange (uptake/excretions) compounds that is produced from the individual members of the community that is not on media can be considered as produced from community members.

## Can I get a list of metabolites coming out from each of the microorganisms?

I would refer to [this narrative](https://narrative.kbase.us/narrative/13838) which provide an example code-cell that printout the exchanges of the community. As an alternative approach/or a workaround, you would be able to check the exchange reactions under reactions tab of the FBA results output. If you search for reactions with \[e0], that will filter all the transporter reactions, were you would see the compounds that excreted out into extracellular compartment or or up-taken in to cytosol compartment of the members in the community.

## I am using KBase to build metabolic models of bacteria. I have questions about the media. Is the concentration of the compounds used in the KBase metabolic model and FBA app? What should I do if there is no production of biomass on a media where the bacteria is supposed to grow in?

Flux Balance Analysis (FBA) is a mathematical approach for analyzing the flow of metabolites through a metabolic network, typically used for predicting growth rates of organisms or the production rates of metabolites under a given set of conditions. The reason why the concentration of compounds in the media does not impact an FBA simulation is primarily due to the assumptions and nature of the model itself. Here's a closer look at why:

1. **Steady-State Assumption**: FBA operates under the assumption that the system is in a steady state, meaning that the concentration of internal metabolites does not change over time. This assumption simplifies the model by focusing only on the fluxes through the metabolic network, rather than the concentrations of the metabolites themselves.
2. **Linear Optimization**: FBA uses linear optimization to predict the flux distribution within the metabolic network that maximizes (or minimizes) a specific objective function, usually related to growth or production of a certain metabolite. Since the model is concerned with fluxes (rates of reactions) rather than metabolite concentrations, the initial concentrations of compounds in the media do not directly affect the optimization process.
3. **Concentration Independence**: The basic formulation of FBA does not include kinetic parameters or concentration-dependent reaction rates. It is based on stoichiometry and constraints on reaction fluxes, such as thermodynamic feasibility and capacity limits of enzymes, but not on the concentrations of substrates or products in the reactions. This simplification allows FBA to be applied without detailed kinetic data, which is often unavailable for complex networks.
4. **Focus on Capabilities of the Network**: The goal of FBA is to understand the capabilities of the metabolic network under a given set of conditions, such as the maximum possible growth rate or the maximum production rate of a metabolite. These capabilities are determined by the network structure and the constraints applied to it, not by the external concentrations of compounds.

However, it's important to note that while basic FBA does not consider compound concentrations, extensions of FBA, such as dynamic FBA (dFBA), can take into account changes in metabolite concentrations over time by integrating FBA with differential equations that describe the dynamics of external metabolite concentrations. These advanced models can provide a more detailed and dynamic understanding of metabolic systems but at the cost of increased computational complexity and data requirements.

The reason that field is available currently in the media object in KBase is for descriptive purposes. While it doesn't impact the FBA simulation, it provides valuable information about your experimental workflow that can be helpful to other researchers you share a narrative with or if you make your narrative public. Also if in the future we implement something like dFBA, it will be easier to just re-run the analysis with the existing media objects without having to edit media objects to add concentration information.

#### We welcome your [feedback](https://www.kbase.us/support) on the information provided in this FAQ. If you have a question about the metabolic modeling tools and data in KBase that is not answered here, or the answer provided didn’t work for you, please [contact us](https://www.kbase.us/support).


# Community Developed Workflows and Tools

Apps and features added by the KBase Community.

Many of the apps you find in KBase were developed in partnership with community developers. You can find their demo Narratives and recorded webinars here.

These apps were developed through partnerships with Department of Energy Science Focus Areas (SFAs). SFAs are collaborative research programs across Department of Energy laboratories and collaborators at other institutions to engage in coordinated, high-quality research.&#x20;

You can read more about the SFAs in general [here](https://genomicscience.energy.gov/sfas/) and more about the SFAs working with KBase [here](https://www.kbase.us/research/user-working-groups/).

1. [Functional Annotation Tools (µBiospheres SFA)](/community-workflows-tools/functional-annotation)
2. [Viral Tools (Microbes Persist SFA)](/community-workflows-tools/viral)
3. [Taxonomy Tools (Bacterial-Fungal Interactions SFA)](/community-workflows-tools/taxonomy)
4. [Functional and Taxonomic Profiling of MAGs (ENIGMA SFA)](/community-workflows-tools/functional-taxonomic-profiling)
5. [Random Walk with Restart toolkit (Exascale Networks)](/community-workflows-tools/rwroolkit)


# Functional Annotation

These apps allow users to import and utilize external genome annotations that can be combined with annotations created by KBase apps such as DRAM, Prokka, or RASTtk.

## Apps

* [Import Annotations From Staging:](https://kbase.us/applist/apps/MergeMetabolicAnnotations/import_annotations/release) Enables upload of third-party annotations like EC and KEGG into KBase.&#x20;
* [Bulk Import Annotations From Staging](https://kbase.us/applist/apps/MergeMetabolicAnnotations/import_bulk_annotations/release): Enables upload of third-party annotations to add to an existing genome into KBase in bulk.
* [Compare Metabolic Annotations](https://kbase.us/applist/apps/MergeMetabolicAnnotations/compare_metabolic_annotations/release): Compare and contrast metabolic annotations from multiple sources.&#x20;
* [Merge Metabolic Annotations](https://kbase.us/applist/apps/MergeMetabolicAnnotations/merge_metabolic_annotations/release): Generate a consensus annotation by combining multiple annotation sources.&#x20;

## Inputs

These apps take [Genomes](/data/upload-download-guide/genome) with annotations as input. See [Assembling & Annotation Microbial Genomes](/workflows/assembly-annotation) for more on annotating genomes.

## Narratives

* &#x20;[Importing annotation sources for metabolic modeling](https://narrative.kbase.us/narrative/66117) – provides an example using a *Streptomyces coelicolor* strain A3(2) and annotations created in KBase and 3rd party tools.
  * Created and made available by Patrik D'haeseleer of Lawrence Livermore National Laboratory.&#x20;
* [Microbial Genomics in KBase: Drafting Isolate Genomes](https://narrative.kbase.us/narrative/83666) and [Microbial Metabolic Model Reconstruction and Analysis](https://narrative.kbase.us/narrative/18302) – provide guidance through related workflows.&#x20;

## Tutorial Webinar

{% embed url="<https://www.youtube.com/watch?v=p6rQS7KgURw>" %}
Tutorial webinar presented by Patrik D'haeseleer of Lawrence Livermore National Laboratory
{% endembed %}

## SFA Collaboration Page&#x20;

These apps were developed as part of the Biofuels and Bioenergy SFA. To learn more about this collaboration, see: <https://www.kbase.us/research/stuart-sfa/>.


# Functional and Taxonomic Profiling of MAGs

These tools allow users to examine the genetic potential of a community  and investigate the strain frequencies of subpopulations in those communities.

## Apps

* [Meta-Decoder Call Variants](https://narrative.kbase.us/#catalog/modules/kb_meta_decoder): Identify polymorphisms within microbiome read sequences against reference genomes.&#x20;
  * [Map Reads to Reference Sequence](https://narrative.kbase.us/#catalog/apps/kb_meta_decoder/map_reads_to_reference/beta): Map reads to a reference genome.
  * [Call Microbial SNPs](https://narrative.kbase.us/#catalog/apps/kb_meta_decoder/call_snps/beta): identify single nucleotide polymorphism (SNP) variants based on mapped reads.&#x20;
* [StrainFinder v1](https://narrative.kbase.us/#catalog/apps/kb_StrainFinder/run_StrainFinder_v1/beta): Determine strain genome sequences from sub-populations of allele frequency.&#x20;
  * [strainFinder Repository](https://bitbucket.org/yonatanf/strainfinder/src/master/)
* [FamaProfiling](https://github.com/aekazakov/FamaProfiling): Generate a functional profile of nitrogen cycle genes or universal single-copy markers for metagenomic read libraries and assembled genomes.
  * [Fama Read Profiling](https://kbase.us/applist/apps/FamaProfiling/run_FamaReadProfiling/release)&#x20;
  * [Fama Genome Profiling](https://kbase.us/applist/apps/FamaProfiling/run_FamaGenomeProfiling/release)&#x20;

## Inputs

These apps take metagenome-assembled genomes (MAGs) as [Genome](/data/upload-download-guide/genome) objects input. See [Metagenomic and Community Analysis](/workflows/metagenomic-analysis) for an overview of metagenomic workflows.

## Narratives

* [Groundwater MAGs from Tian et al., 2020](https://narrative.kbase.us/narrative/90581) – provides an example of the workflow using MAGs
  * Tian et al. (2020) *Small and mighty: adaptation of superphylum Patescibacteria to groundwater environment drives their genome simplicity*. Microbiome 8:51. <https://doi.org/10.1186/s40168-020-00825-w>
  * Narrative created and made public by Alexey Kazakov, Lawrence Berkeley National Laboratory
* [Metagenome-Assembled Genome Extraction from a Compost Microbiome Enrichment](https://narrative.kbase.us/narrative/33233) – guidance for creating MAGs.

## Tutorial webinar

{% embed url="<https://www.youtube.com/watch?v=Poa4cW3hXU8>" %}
Tutorial webinar presented by by John-Marc Chandonia (Lawrence Berkeley National Laboratory), Alexey Kazakov (Lawrence Berkeley National Laboratory), and Anni Zhang (Massachusetts Institute of Technology)
{% endembed %}

## SFA Collaboration Page

These apps were developed as part of the Ecosystems and Networks Integrated with Genes and Molecular Assemblies (ENIGMA) Science Focus Area (SFA). To learn more about this collaboration, see: <https://www.kbase.us/research/adams-sfa/>.


# Taxonomy

These classifier tools enable users to assign taxonomy and detect unique genomic signatures.

## Apps

* [GOTTCHA2 (**G**enomic **O**rigin **T**hrough **T**axonomic **Cha**llenge)](https://kbase.us/applist/apps/gottcha2/run_gottcha2/release): Uses a novel taxonomic profiling method for metagenomes that is both gene-independent and signature-based. It significantly reduces false discovery rates (FDR).&#x20;
* [Centrifuge Taxonomic Classifier](https://narrative.kbase.us/#catalog/apps/centrifuge/run_centrifuge/beta): Enables rapid, accurate, and sensitive labeling of metagenomic reads and quantification of species. *(Beta-only)*
* [Kraken2 Taxonomic Sequence Classifier](https://narrative.kbase.us/#catalog/apps/kraken2/run_kraken2/beta): Examines the k-mers within a query sequence and maps k-mers to the lowest common ancestor (LCA) of all genomes in the reference database known to contain a given k-mer. (*Beta-only*)

## Input

These apps take [single- and paired-end reads](/data/upload-download-guide/reads) or sets of reads as input.

## Narratives

* [GOTTCHA2 webinar](https://narrative.kbase.us/narrative/88330) – demonstrates Genomic Origin Through Taxonomic CHAllenge apps using example data.
  * Created and made public by Mark Flynn of Los Alamos National Laboratory.
* [Metagenome-Assembled Genome Extraction from a Compost Microbiome Enrichment](https://narrative.kbase.us/narrative/33233) – for more on how to create metagenome-assembled genomes.&#x20;
* [Kaiju](https://narrative.kbase.us/legacy/catalog/apps/kb_kaiju/run_kaiju/release) – a similar taxonomic classification app.

## Tutorial Webinar

{% embed url="<https://www.youtube.com/watch?v=N-_89NJu9So>" %}
Tutorial webinar presented by Mark Flynn of Los Alamos National Laboratory
{% endembed %}

## SFA Collaboration Page

These apps were developed as part of the Bacterial-Fungal Interactions SFA at LANL. To learn more about this collaboration, see: [https://www.kbase.us/chain-sfa/](https://www.youtube.com/redirect?event=video_description\&redir_token=QUFFLUhqbG8yUEZKNE9FN3ZmbnNVdjIwNE5xUlBESFZnd3xBQ3Jtc0ttU1FPR0U3UU55dmRmUFJWd0piS0NMcE8zNVM5SHBFS1o0SFRBendmeW1TQTNlREdJQVlmNE5KTUhpaWFyR3MtR2w5N2xjckxvRF9fdVVwbzNRYnE2VnJCUWxxZVVJMjNWQlpwMkRoeElvbEFXbUhETQ\&q=https%3A%2F%2Fwww.kbase.us%2Fchain-sfa%2F\&v=N-_89NJu9So).


# Viral

These apps allow users to identify and analyze viral sequences in assemblies within KBase.

## Apps

* [vContact2](https://kbase.us/applist/apps/vConTACT/vcontact/release): perform guilt-by-contig-association automatic classification of viral contigs
* [VirSorter2](https://kbase.us/applist/apps/kb_virsorter2/run_kb_virsorter2/release): Identifies viral sequences from viral and microbial metagenomes
  * VirSorter (legacy)
* [VirMatcher](https://narrative.kbase.us/#catalog/apps/kb_virmatcher/run_kb_virmatcher/release):  Predicts host-virus matches

## Input

The viral workflow begins with [Assemblies](/data/upload-download-guide/assembly), either metagenome or single genomes. VirSorter produces the assemblies that should be used by vContact as input. &#x20;

## Narratives

* [Viral Annotation Pipeline in KBase](<	https://doi.org/10.25982/59912.37/1635154>) – a streamlined use case of a pre-assembled small metagenome and a Cyanophage genome run through VirSorter to identify viral sequences and vContact to classify them.
* [Viral Analysis End-to-End](https://doi.org/10.25982/75811.86/1868934) – a more comprehensive tutorial with a use case from the Global Ocean Virome demonstrating how to go through the entire process starting with raw metagenomic reads.
  * Tutorial Narratives were made available by Ben Bolduc of Ohio State University.&#x20;

## Tutorial Webinar

{% embed url="<https://www.youtube.com/watch?v=WbQeOzdyTbc>" %}
Tutorial webinar presented by by Ben Bolduc (Ohio State University).
{% endembed %}

## SFA Collaboration Page

These apps were developed as part of the Microbes Persist SFA. To learn more about this collaboration, see: <https://www.kbase.us/research/pett-ridge-sfa/>.


# Random Walk with Restart Toolkit

These apps allow users to explore networks of multiplex data in an interactive network visualization using the Random Walk with Restart Toolkit (RWRtools).

## Apps

* [Find Functional Context using Lines of Evidence with RWRtools LOE](https://narrative.kbase.us/#catalog/apps/kb_djornl/run_rwr_loe/): uses Random Walk with Restart (RWR) to rank genes in the network starting from a Feature Set.
* [Find Gene Set Interconnectivity using Cross Validation with RWRtools CV](https://narrative.kbase.us/#catalog/apps/kb_djornl/run_rwr_cv): performs cross validation on a single gene set, finding the RWR rank of the left-out genes.

## Inputs

The input for this app is an annotated [Genome](/data/upload-download-guide/genome).&#x20;

The network data is maintained by KBase behind the scenes and can be found [here](https://github.com/kbaseapps/kb_djornl/blob/main/NETWORKS.md).

## Narratives

* [RWRtools Exascale and Petascale Network Analysis Demo Narrative](https://narrative.kbase.us/narrative/167526) – was created as part of the software publications for the RWRtoolkit (manuscript in preparation).
  * Created and made public by Kyle Sullivan of Oak Ridge National Laboratory
* [Microbial Genomics in KBase: Gene Feature Analysis](https://narrative.kbase.us/narrative/83681) – provides further examples of creating and working with FeatureSets derived from Genomes.

## Tutorial webinar

{% embed url="<https://www.youtube.com/watch?v=Dq5kZ_Sacy0>" %}
Tutorial webinar presented by Kyle Sullivan, Oak Ridge National Laboratory
{% endembed %}

## SFA Collaboration page

These apps were developed as part of the Exascale Network project led by Dan Jacobson. To learn more about this collaboration, see: <https://www.kbase.us/research/user-working-groups/#datascience>.


# Troubleshooting

This section provides guidance on troubleshooting and issue reporting. We greatly appreciate your feedback as we strive to make KBase a resource for the community. Please don't hesitate to contact us if you encounter a bug or want to suggest a new feature using the rel listed below.&#x20;

## Contents

1. [Experiencing problems with the Narrative User Interface](/troubleshooting/narrative)
2. [The Help Board](/troubleshooting/support)
3. [Reporting issues to the Help Board](/troubleshooting/report)
4. [Finding an error in the Job Log](/troubleshooting/job-errors/common/job-log)
5. [Common error messages and their meanings](/troubleshooting/job-errors)


# Problems with the User Interface

### If you experience problems with KBase’s Narrative user interface (e.g., data, apps, or viewers won’t load or the interface appears to be frozen), here are a few things to try that may help:

* See our list of [recommended web browsers](/getting-started/browsers#supported-browsers) to make sure your browser is supported by KBase.
* Try a hard reload of your Narrative (shift-reload in your browser window).
* Log out of KBase and log in again.
* Clear your browser cache or open the Narrative Interface in an “incognito” or “private” browser window.
* Select the “Shutdown and Restart” option in the menu at the top left of the Narrative Interface. A popup will ask you to confirm that you want to restart the Narrative Interface and remind you to save your latest changes.

### What should I do if my Narrative freezes or won’t load properly?

* See our list of [recommended browsers](/getting-started/browsers#supported-browsers) to ensure yours is supported by KBase.
* Try a hard reload (shift-reload in your browser window).
* If you’ve used earlier versions of KBase, they may be lurking in your cache and messing things up. Try either clearing your cache or [launching KBase](https://narrative.kbase.us/) from an “incognito” or “anonymous” window (which doesn’t use your cache).
* Shut down and restart your Narrative by selecting this option from the top left menu of the Narrative Interface.


# Help Board

## Contact Us

The [KBase Help Board](< https://kbase-jira.atlassian.net/>) is an interactive way for users to report bugs, ask questions, or suggest new features, using an issue tracking system called Jira.

Using the [Help Board](< https://kbase-jira.atlassian.net/>), you’ll be able to:

* [Submit bug reports, questions and suggestions](/troubleshooting/report) without revealing your email address
* Engage in a two-way dialog with KBase staff as they track down and resolve your issue
* Search the issues submitted by other users. If someone else has already reported your problem or suggested the same new feature, you can vote for that issue or add a “Me too!” comment.
* Increase your productivity in KBase by seeing other users’ suggestions and workarounds

## **Access the Help Board**

1. To get into Jira to use the KBase Help Board, go to <https://kbase-jira.atlassian.net/>.&#x20;
2. Log in with your Google account or click “Sign up for an account.”

See more on how to report issues to the Help Board to find out how to [submit your bug reports, questions and suggestions](< https://kbase-jira.atlassian.net/>) to us.


# How to Report Issues

To report issues, suggest new features, or get help using KBase tools, please go to our [Help Board](< https://kbase-jira.atlassian.net/>) and create a ticket.&#x20;

![](/files/-MB07md4vjclA4TXF3jM)

## Create a Ticket

On the [KBase Help Board](https://kbase-jira.atlassian.net/), click the blue **Create** button in the menu. A window pops up to gather information about your issue. *Required fields are marked with a red \**; the other fields are optional.&#x20;

1. Choose the **issue type** (*Bug,* *Question*, or *New Feature*)
2. Enter a **short summary**
3. Enter a longer **description** with more details
4. Enter the **URL of the affected Narrative** (remember to [share the Narrative](/getting-started/narrative/share) with "kbasehelp")
5. Include the specific **Job ID** found within the [Job Log](/troubleshooting/job-errors/common/job-log#job-browser)
6. Select the best fit **Category**
7. Include you **KBase Username**
8. The **Environment** field can be used to let us know which operating system you’re using (e.g., Mac, Windows XP, etc.) and which web browser (e.g., IE 10, Chrome, etc.)
9. **Attach files,** such as screenshots or small data files that demonstrate the problem
10. Include the **Incident Start and End Times** if applicable&#x20;
11. Click the blue **Create** button to submit your ticket

Remember, the more information provided about what went wrong, the more efficiently our team can reproduce and resolve the problem.


# Job Errors and Their Meanings

Some job errors give an explanation of the problem and others can be much more cryptic. Here are more common job error messages found within the [Job Log](/troubleshooting/job-errors/common/job-log) along with the context of the App or job-associated, the meaning of the error, and steps to take toward a solution.&#x20;

If you have an error message and are looking to find a solution, start with the Common Job Errors or select the specific app category page to use the browser's search function to look for the error message and solutions.&#x20;

1. [Common Job Errors](/troubleshooting/job-errors/common) – occur across all apps
2. [Import Job Errors](/troubleshooting/job-errors/import) – specific to importing data
3. [Assembly App Errors](/troubleshooting/job-errors/assembly) – messages specific to Assembly apps
4. [Annotation App Errors](/troubleshooting/job-errors/annotation) – messages specific to Annotation apps
5. [Functional Genomics App Errors](/troubleshooting/job-errors/functional-genomics) – messages specific to apps relevant to Functional Genomics analysis
6. [Modeling App Errors](/troubleshooting/job-errors/modeling) – messages specific to Modeling apps


# Common Job Errors

This is a list of common job error messages found within the Job Log for all KBase Apps, what they mean and how to go about fixing them or if a job ticket needs to be submitted.&#x20;

You can always submit a ticket for help, questions, or follow-up to the [KBase Help Board](https://kbase-jira.atlassian.net/jira/your-work).&#x20;

## **Types of Errors**

* **UE**=User fixable or possible user error
* **KE**=KBase error
* **UK**=Unknown or multiple possibles causes
* **TE**=Temporary system problem

{% hint style="info" %}
Use your Browser's search tool to paste in your error message to locate your error message and next steps.&#x20;
{% endhint %}

## Common Job Error Messages

#### `(2, No such file or directory)`&#x20;

UK: The app did not produce any output or it produced output but none passed the filters. There may also be app-specific reasons described below.&#x20;

*Check the logs for more information and adjust filters as needed.*&#x20;

#### `Unable to build output viewer`&#x20;

UK: The job ran for more than 7 days and crashed. There is no output. Dependent on the reason for the crash, rerunning the job may work.&#x20;

*Resubmit job.*&#x20;

#### `Output file is not found, exit code 123`&#x20;

UK: No space left on device. The job may be too big for KBase in the current configuration.&#x20;

*Resubmit job to make certain the issue isn’t a conflict with other jobs. If this doesn’t work, submit a ticket to the* [*Help Board*](https://kbase-jira.atlassian.net/projects/PUBLIC/issues)*.*&#x20;

#### `Output file is not found, exit code is 137`&#x20;

UK: Cause unknown. Job probably cancelled by another process but not the user.&#x20;

*Resubmit job.*

#### `Job was cancelled as it ran overMax allotted time (604800000) milliseconds (10080) minutes`&#x20;

UE: Job ran for more than 7 days and finished cleanly. Likely the job is too big.&#x20;

*Resubmit job, dependent on the reason it took so long.*&#x20;

#### `'listener timeout after waiting for [600000] ms'`&#x20;

TE: The utility ElasticSearch went down. The fix is a manual process and may not get fixed during nights, weekends, and holidays.&#x20;

*Submit a ticket the the* [*Help Board*](https://kbase-jira.atlassian.net/projects/PUBLIC/issues)*. Resubmitting the job every couple of hours until it runs may also work.*&#x20;

#### `kafka`&#x20;

TE: Any messages with the word kafka in the error. Something went wrong with the system.&#x20;

*Resubmit job. Report the issue to the* [*Help Board*](https://kbase-jira.atlassian.net/projects/PUBLIC/issues) *if resubmitting doesn’t fix the problem.*&#x20;

#### `ProtocolError, Connection aborted., BadStatusLine`&#x20;

UK: Something went wrong with the reporting and cleanup at the end of the job. Intermittent error. The data is fine, but there will be no report at the end. If an object was created, clicking on it in the data panel will create a viewer for the object, which is likely missing.&#x20;

*Resubmit job if you need the end report.*

#### `User XXXX may not read workspace nnnnn`&#x20;

UE: The app requires data owned by another user and you do not have access. In another scenario, either you or the user have deleted the original file.&#x20;

*Ask the user for access to the needed file or re-import file if deleted.*&#x20;

#### `GET` [`http://elasticsearch.kbase.us:9200/kbaseprod.taxon_*,-*_sub/data/_search`](http://elasticsearch.kbase.us:9200/kbaseprod.taxon_*,-*_sub/data/_search)`: HTTP/1.1 503 Service Unavailable`&#x20;

TE: A component in KBase needed to be rebooted and will recover in a few minutes.&#x20;

*Try resubmitting the job in 10-15 minutes.*&#x20;

#### `No such container:...`

TE: A known temporary error.&#x20;

*Resubmit job.*&#x20;

#### `Token validation failed: Too many open files'`&#x20;

TE: A known temporary error .

*Try resubmitting the job in 10-15 minutes.*&#x20;

#### `502 Server Error: Bad Gateway….`&#x20;

TE: A known temporary error.

*Resubmit job.*&#x20;

#### `Gee whiz, I sure am sorry, but an error occurred. Gosh!...... Object 19 cannot be accessed`

UE: User’s browser cache is retaining old information. In this case, it is 'Object 19'. One symptom is that the user gets the error but the narrative looks fine to others.&#x20;

*Try the following:*

* *Reload the page*&#x20;
* *Close and reopen your narrative*&#x20;
* *Log out and log back in again*
* *Clear your cache*

*If you find something that works, no need to try the others.*

#### `Details: 500 globus_xio: ICE negotiation failed.`

UE: The error generally indicate a firewall issue.&#x20;

*Please refer to this question on the Globus forum for a similar question with its solution. If you are working from home, you can view the Globus documentation for configuring the firewall. If you are on an institution network, you'll probably have to request an exception from IT.*


# The Job Log

The Job Log contains run information from apps and other jobs, including important error information.

The Job Log can be found in two places: an app's *Job Status* tab and the **Job Browser** (in the sidebar Menu).

### **Job Status tab**

The Job Log starts at the bottom with the error highlighted in pink. There may be useful information located just above the pink section.&#x20;

![](/files/-MN0O2FEpKMZgPKR1M7X)

### **Job Browser**

Find the line with the job of interest and click on the icon for the piece of paper. This will open up the log, starting at the top.&#x20;

<div align="center"><img src="/files/-M7teuVNgSpaOgIJoyax" alt=""></div>

In the Job Browser, jobs not associated to an app i.e., export, are included within the job list. The Job ID and worker node are easily located and the log can be scrolled through within the pop-out. Job Logs can be downloaded in CSV, TSV, JSON, and TEXT formats using the download (downward-facing arrow into tray) icon.&#x20;

![](/files/-M7xf8Dl9bO1zZm03vfC)

{% hint style="info" %}
Even if you find an explanation for your error in the [list of common errors](/troubleshooting/job-errors), it is okay to [submit a Help Board ticket](/troubleshooting/support) with follow-up questions.
{% endhint %}

### Helpful Points

* If the output object is created but the job ends in an error, assume that the error didn’t affect the data. Errors sometimes occur in the cleanup and reporting at the end while not affecting the output data.
* Jobs *do not run more than 7 days*. If the narrative and/or the Job Browser say that a job is still running after 7 days, assume that it is dead. Resubmitting may solve the problem, but it is possible the job might be too big for KBase.&#x20;
* Assemblers currently have an upper limit of between 180,263,840 paired reads and 240,351,788 reads depending on complexity. If the job has been run twice, exceeded the 7 day limit, and your data is in this size range, it may be too big for KBase at this time.
* Bowtie appears to have an upper limit between 3.5E+10 and 4.0E+10 bases due to space limitations.
* If asked for a job ID, it is a 24-digit hexadecimal number that looks similar to this: 5e1e3234e4b0fb2e6517240a. It is in the first line of the job log.
* If asked for the node where the job ran, it is on the first line of the job log and looks similar to `Running on chicagoawe163` or `Running on uploadworker` or `Running on uploadworker-prod-dtn2`
* Any file uploaded from a Windows machine might have  DOS-style carriage-return line files along with new-lines. This has been known to cause problems with importing FASTQ files  and may affect other files as well.


# Import Job Errors

This is a list of error messages found within the Job Log for Import Jobs, what they mean and how to go about fixing them or if a job ticket needs to be submitted.

## **Types of Import Job Errors**

* **UE**=User fixable or possible user error
* **KE**=KBase error

{% hint style="info" %}
Use your Browser's search tool to paste in the error to locate possible fixes or next steps.&#x20;
{% endhint %}

You can always submit a ticket for help, questions, or follow-up to the [KBase Help Board](https://kbase-jira.atlassian.net/jira/your-work).&#x20;

## Common Import Job Error Messages

#### `Cannot connect to URL: ftp://ftp.imicrobe.us. ...`&#x20;

UE: The provided URL cannot be accessed from within KBase.&#x20;

*Recheck the URL and permissions. Then try resubmitting the job.*&#x20;

#### `Invalid FTP Link: ...`&#x20;

UE: The provided URL cannot be accessed from within KBase. Perhaps the option for ‘Direct’ download should be specified instead of ‘FTP’ (e.g., when downloading from the SRA)&#x20;

*Recheck the URL and permissions. Then try resubmitting the job.*&#x20;

#### `Invalid Google Drive Link ...`&#x20;

UE: The provided URL cannot be accessed from within KBase.&#x20;

*Recheck the URL and its permission. Then try resubmitting the job.*&#x20;

## **Bulk or Batch Imports**

#### `(2, No such file or directory)`&#x20;

UE: The file is not in the Staging Area.

*Verify that the name is correct and upload is complete. Then try resubmitting the job.*&#x20;

## [**Import FASTQ/SRA File as Reads from Staging Area**](https://narrative.kbase.us/#catalog/apps/kb_uploadmethods/import_fastq_sra_as_reads_from_staging)

#### `(2, No such file or directory)`&#x20;

KE: The fastqdump ran but the file names are not the expected names.&#x20;

*Use the long workaround* [*here*](https://kbase-jira.atlassian.net/browse/PUBLIC-946)*.*

#### `SRA input file type selected. But missing SRA file`&#x20;

U&#x45;**:** The format of the file is not recognized.&#x20;

*Recheck the file and try resubmitting the job.*&#x20;

#### `Invalid FASTQ file ...`&#x20;

KE/UE: Sometimes the user has specified the file name wrong. It can also happen because the importer has problems with file names that end in ".1"&#x20;

*Use the long workaround* [*here*](https://kbase-jira.atlassian.net/browse/PUBLIC-946)*.*

#### `Error running command: /kb/deployment/bin/fastq-dump ...`&#x20;

UE: The file does not appear to be in the expected SRA format.&#x20;

*Recheck the file and try resubmitting the job.*&#x20;

#### `Error running command:pigz ...`&#x20;

UE: The file could not be unzipped by KBase and most likely couldn’t be unzipped by the user either.&#x20;

*Verify the file is can be unzipped locally.*&#x20;

#### `Both SRA and FASTQ/FASTA file given.`&#x20;

UE: The inputs should be either all fastq/a or all SRA.&#x20;

*Modify the inputs, then try resubmitting the job.*&#x20;

#### `Same file [XXX.XXXX.gz] is used for forward and reverse. Please select different files and try again.`&#x20;

UE: There are names for both a forward and reverse strand and they are identical.&#x20;

*A Single-end read library only needs one name. A Paired-end read library needs two files with different names.*&#x20;

#### `File /kb/XXX.fasta is not a FASTQ file`&#x20;

UE: Either the file is not in fastq format or the file extension is not recognized.

*Recheck that the file is in the right format. Change the extension to .fastq if needed, then try resubmitting the job.*&#x20;

#### `Invalid FASTQ file`&#x20;

UE: Possible issues

* The fastq file includes one or more sequences that are less than 10 bases. Short reads are a problem for some tools.
* The fastq file doesn't have the right number of lines. For example, the lines in a single-end file needs to be a multiple of four and interleaved paired-end library should be a multiple of eight.
* The options haven't been selected correctly. For example, using an interleaved fastq file but failing to check the Interleaved box. *The documentation on* [*FASTQ/SRA Reads*](/data/upload-download-guide/reads) *may be helpful.*
* The file might not have the right filename to be recognized.
  * The file is an SRA file and not FASTQ.&#x20;
* DOS-style carriage-return line files along with new-lines. Our fasta validation doesn't handle this properly. *To remove the carriage return characters use this* [*unix command:*](https://support.nesi.org.nz/hc/en-gb/articles/218032857-Converting-from-Windows-style-to-UNIX-style-line-endings) ***`tr -d '\015' < 1.fastq >cleaned_1.fastq`***

#### `Reading FASTQ record failed - non-blank lines are not a multiple of four.`

UE: The number of lines in the FASTQ file are not a multiple of four.

*Recheck the file and try resubmitting the job.*&#x20;

#### `Interleave failed - reads files do not have an equal number of records….`

UE: Something went wrong trying to interleave the Paired-end files.

*Recheck the line count of the files. Hidden carriage returns or linefeeds in the file could contribute to the problem.*

#### `Deinterleave failed - line count is not divisible by 8`

UE: The interleaved file does not appear to be the correct format.

*Recheck the file and try resubmitting the job.*&#x20;

#### `Object 1: Illegal character in object name`&#x20;

UE: The name of the output reads object can’t have spaces or special characters.

*Rename the output file and then try resubmitting the job.*&#x20;

## [Import FASTA File as Assembly from Staging Area](https://narrative.kbase.us/#catalog/apps/kb_uploadmethods/import_fasta_as_assembly_from_staging)

#### `There are no contigs to save, thus there is no valid assembly.`

U&#x45;*:* There are no contigs that passed the minimum contig size.

*Adjust the minimum contig size or other optional parameters. Then try resubmitting the job.*&#x20;

#### `The FASTA header key XXX appears more than once in the file`

UE: The FASTA header lines may not be unique.

*Recheck the format of the header lines and* *try resubmitting the job.*&#x20;

#### `This FASTA file has non nucleic acid characters`

UE: The file appears to be proteins or special characters instead of DNA.

*Recheck the file contents, and then* *try resubmitting the job.*&#x20;

#### `This FASTA file may have amino acids in it instead of the required nucleotides.`

UE: The file appears to be proteins instead of DNA.

*Recheck the file contents, and then* *try resubmitting the job.*&#x20;

#### `FASTQ/FASTA input file type selected. But missing FASTQ/FASTA file`

UE: Selected file does not match the import selected.

*Select a valid combination and try resubmitting the job.*&#x20;

#### `(\utf-8\, b\PK\x03\x04\x14\x00\x08…….`

UE: Attempt to import a zip file with multiple files as a single data object.

*Run the App '* [*Unpack a Compressed File in Staging Area*](https://narrative.kbase.us/#catalog/apps/kb_uploadmethods/unpack_staging_file)*' on the file and retry resubmitting the job.*&#x20;

## [Import GenBank File as Genome from Staging Area](https://narrative.kbase.us/#catalog/apps/kb_uploadmethods/import_genbank_as_genome_from_staging)

#### **`Duplicate gene ID: XXXX_xxxx`**

UE: Gene IDs within the input file are not unique.

*Edit gene IDs and try resubmitting the job.*&#x20;

#### **`The input directory does not have any files with one of the following extensions .gbff,.gbk,.gb,.genbank,.dat,.gbf`**

UE: The app only recognizes the listed file extensions as valid GenBank files.

*Change the file extension and try resubmitting the job.*&#x20;

#### **`XXX is not a valid KBase taxon ID`**

UE: The Taxonomy ID in the advanced parameters is optional and needs to be an integer when specified. The user provided the text string ‘XXX’.

*Use an integer taxon ID or leave it blank. The information will be picked up from the GenBank file or from the scientific name.*

## [Import GFF3/FASTA file as Genome from Staging Area](https://narrative.kbase.us/#catalog/apps/kb_uploadmethods/import_gff_fasta_as_genome_from_staging)

#### **`Every feature sequence id must match a fasta sequence id`**

UE: The IDs in the ‘sequence source’ lines must match the header lines in the FASTA file in the GFF format.&#x20;

*Correct the GFF format and try resubmitting the job.* &#x20;

#### **`unable to parse >.....`**

UE: The file may not be in GFF format.&#x20;

*Recheck the file format and try resubmitting the job.*&#x20;

#### `Features must be completely contained within the Contig in the Fasta file.`

UE: The coordinates for the feature are outside the bounds of the contig.

&#x20;*Recheck the file where indicated and try resubmitting the job. In rare instances, the GFF file contains a feature that wraps around the 0 position and the coordinates look like the feature goes off the end of the sequence. The options are to 1) remove the feature from the GFF file, 2) edit the feature so that it is in two parts, or 3) find a GenBank formatted version of the file and resubmit.*&#x20;

## [Import Media file (TSV/Excel) from Staging Area](https://narrative.kbase.us/#catalog/apps/kb_uploadmethods/import_tsv_excel_as_media_from_staging)

#### **`Starch.csv" is not a valid EXCEL nor TSV file`**

UE: File format is not recognized.&#x20;

*Recheck the file format and try resubmitting the job.*&#x20;

## [Import TSV/XLS/SBML File as an FBAModel from Staging Area](https://narrative.kbase.us/#catalog/apps/kb_uploadmethods/import_file_as_fba_model_from_staging)

#### **`data file 4.xml either does not use commas or tabs as a separator`**

UE: File format is not recognized.&#x20;

*Recheck the file format and try resubmitting the job.*&#x20;

#### **`No object with name _Nostoc_azollae__0708`**

UE: The genome does not exist in the narrative&#x20;

*Fix the genome name and try resubmitting the job.*&#x20;

## [**Unpack a Compressed File in Staging Area**](https://narrative.kbase.us/#catalog/apps/kb_uploadmethods/unpack_staging_file)&#x20;

#### `Error running command:pigz ...`&#x20;

UE: The file could not be unzipped by KBase and most likely couldn’t be unzipped by the user either.&#x20;

*Verify the file can be unzipped locally.*&#x20;


# Assembly App Errors

This is a list of error messages found within the Job Log for Assembly App Jobs, what they mean and how to go about fixing them or if a job ticket needs to be submitted.

## **Types of Errors**

* **UE**=User fixable or possible user error
* **KE**=KBase error
* **UK**=Unknown or multiple possibles causes
* **TE**=Temporary system problem

{% hint style="info" %}
Use your Browser's search tool to paste in your error message to locate your error message, information about the error and next steps.&#x20;
{% endhint %}

You can always submit a ticket for help, questions, or follow-up to the [KBase Help Board](https://kbase-jira.atlassian.net/jira/your-work).&#x20;

## Common errors across SPAdes Assembly Apps

### [Assemble Reads with SPAdes](https://narrative.kbase.us/#catalog/apps/kb_SPAdes/run_SPAdes); [Assemble Reads with HybridSPAdes](https://narrative.kbase.us/#catalog/apps/kb_SPAdes/run_hybridSPAdes/release); [Assemble Reads with metaSPAdes](https://narrative.kbase.us/#catalog/apps/kb_SPAdes/run_metaSPAdes/release); [Filter Assembled Contigs by Length](https://narrative.kbase.us/#catalog/apps/kb_assembly_compare/run_filter_contigs_by_length)

#### `There are no contigs to save, thus there is no valid assembly.`&#x20;

UE: There are no contigs that passed the minimum contig size.&#x20;

*Adjust the minimum contig size or other optional parameters. Then try resubmitting the job.*&#x20;

#### `Error running SPAdes, return code: 1`&#x20;

UE: Potential explanations&#x20;

* The coverage of your input is so uneven that everything is disconnected.&#x20;
* The reads contain too many k-mers to fit into available memory.&#x20;
* Incomplete write! Reason: No space left on device.&#x20;
* Coverage not uniform hybridSPAdes ended abnormally&#x20;
* Failed to align paired reads left paired reads is not equal to right paired reads cannot specify any data types except a single paired-end library (optionally accompanied by a single library of TSLR-contigs, or PacBio reads, or Nanopore reads) in metaSPAdes mode&#x20;
* The program was terminated by segmentation fault&#x20;
* Too many erroneous kmers, the estimates might be unreliable

*Look for an explanation from SPAdes in the log. If possible, correct the error and try resubmitting the job.*&#x20;

#### `Deinterleave failed - line count is not divisible by 8`&#x20;

UE: The interleaved file does not appear to be the correct format.&#x20;

*Recheck that the files are labeled properly and try resubmitting the job.*&#x20;

#### `Reads object XXX is marked as containing metagenomic data but the assembly method was not specified as metagenomic`&#x20;

UE: The input object and the selected parameters disagree&#x20;

*Change the data input or the app parameters and try resubmitting the job.*&#x20;

#### `Plasmid assembly requires that one and only one library as input.`&#x20;

UE: There is more than one input library.&#x20;

*Change the input to be just one library, change the library source, or merge the libraries and try resubmitting the job.*&#x20;

#### `Invalid type for object 27459/32/1…..`&#x20;

UE: The user is allowed to enter an Assembly for the long reads (like in the MaSuRCA app) but this results in an error.&#x20;

*Ensure objects correspond with the object types listed and try resubmitting the job.*&#x20;

#### `QUAST reported an error, return code was 4`&#x20;

UE: None of the assembly files contains correct contigs (none greater than the minimum contig filter).&#x20;

*Provide different files or decrease --min-contig threshold. Then try resubmitting the job.*&#x20;

## [Velvet Assembler](https://narrative.kbase.us/#catalog/apps/Velvet/run_velvet)

#### `Error on Object #1: Illegal character in object name`

UE: Velvet did not assemble any contigs longer than the minimum length.&#x20;

*Change the configuration and try resubmitting the job.*&#x20;

## [Assemble with HipMer](https://narrative.kbase.us/#catalog/apps/hipmer/run_hipmer_hpc)

#### `'list index out of range'`&#x20;

UE: There is no input data&#x20;

*Add* [*data*](/getting-started/narrative/add-data) *to the app and try resubmitting the job.*&#x20;

#### `Error in HipMER execution`&#x20;

UE: Possible causes -&#x20;

* The data was a metagenome, but the metagenome flag was not enabled&#x20;
* The  metagenome flag was enabled but the data was not a metagenome&#x20;
* Exceeded time limit at NERSC
* Other data/parameter combinations that result in no returned assembly

*Change the configuration and try resubmitting the job.*&#x20;

#### `Using input parameters, you have filtered all contigs from the HipMer assembly. Decrease the minimum contig size and try again`

UE: The minimum contig length is too large.&#x20;

*Lower the minimum contig length in the parameters and try resubmitting the job.*&#x20;

#### `(2, No such file or directory)`&#x20;

TE: Hipmer was in the queue too long and did not run.&#x20;

*Try resubmitting the job*.&#x20;

## [MaSuRCA Assembler](https://narrative.kbase.us/#catalog/apps/kb_MaSuRCA/run_masurca_assembler)

#### `Error running command: XXX/config.txt Exit Code: 1`&#x20;

UE: invalid file for PACBIO invalid file for NANOPORE.

*Recheck the file and try resubmitting the job.*&#x20;

## [Assemble Transcripts using StringTie](https://narrative.kbase.us/#catalog/apps/kb_stringtie/run_stringtie/)

#### `Genome at 48203/40/2 does not have reference to the assembly object`&#x20;

UE: A previous step (e.g., HISAT) used an assembly instead of a genome as input.&#x20;

*Go back to the previous step and use a genome as input.*&#x20;

#### `coercing to unicode: need string or buffer`&#x20;

KE: genome related issue with some of the NCBI RefSeq genomes.&#x20;

*Workaround: import the GFF-format genome with correct feature annotations and try resubmitting the job. See* [*https://kbase-jira.atlassian.net/browse/PUBLIC-969*](https://kbase-jira.atlassian.net/browse/PUBLIC-969)

## [**Merge Reads Libraries**](https://narrative.kbase.us/#catalog/apps/kb_ReadsUtilities/KButil_Merge_MultipleReadsLibs_to_OneLibrary/)

#### `incompatible read library types in ReadsSet`&#x20;

U&#x45;*:* Every library in the input to merge must be of the same type, either all Paired-End or Single-End.&#x20;

*Change the inputs and try resubmitting the job.*

#### **`[Errno 28] No space left on device`**

UE/KE: Most likely case: The combined size of the libraries exceeds the current KBase hardware. Efforts are being made to increase capacity and to provide a cleaner more informative message. Workarounds should be considered.

*The  workarounds include:* &#x20;

1. *Reduce the number of libraries in the merge*
2. *Run Kaiju to determine which samples have similar profiles. Then try the co-assemblies pairwise using two-at-a-time based on which profiles are most similar.*&#x20;

&#x20;&#x20;


# Annotation App Errors

This is a list of error messages found within the Job Log for Annotation App Jobs, what they mean and how to go about fixing them or if a job ticket needs to be submitted.

## **Types of Errors**

* **UE**=User fixable or possible user error
* **KE**=KBase error
* **UK**=Unknown or multiple possibles causes
* **TE**=Temporary system problem

{% hint style="info" %}
Use your Browser's search tool to paste in your error message to locate your error message, information about the error and next steps.&#x20;
{% endhint %}

You can always submit a ticket for help, questions, or follow-up to the [KBase Help Board](https://kbase-jira.atlassian.net/).&#x20;

## **Common Errors across Annotation Apps**

#### RAST: [Annotate Microbial Genome](https://narrative.kbase.us/#catalog/apps/RAST_SDK/reannotate_microbial_genome/) / [Annotate Multiple Microbial Genomes](https://narrative.kbase.us/#catalog/apps/RAST_SDK/reannotate_microbial_genomes/) / [Annotate Microbial Assembly](https://narrative.kbase.us/#catalog/apps/RAST_SDK/annotate_contigset/) **/** [Annotate Multiple Microbial Assemblies](https://narrative.kbase.us/#catalog/apps/RAST_SDK/annotate_contigsets/)&#x20;

#### Prokka:  [Annotate Assembly and Re-annotate Genomes with Prokka](https://narrative.kbase.us/#catalog/apps/ProkkaAnnotation/annotate_contigs/)

#### `(2, No such file or directory)`&#x20;

UE: One or more contigs had header lines longer than 37 characters.&#x20;

*Edit the fasta file, upload again, and try resubmitting the job.*

#### `Error invoking method call_features_rRNA_SEED`&#x20;

UE: One or more of the FASTA header lines are extremely long.&#x20;

*Shorten them, re-upload, and try resubmitting the job.* *.*&#x20;

#### `Object #1, 0ae7adb1-6351-4d26-a526-cc09c15e46ee.report has invalid reference: ….`

UE: There are no input datasets listed.&#x20;

*Supply the dataset(s) and try resubmitting the job.*

#### &#x20;`too many contigs`&#x20;

UE: RAST has a limit of 10,000 contigs. Prokka has a limit of 30,000 contigs.

*Divide the file or bin the contigs before running RAST again by resubmitting the job.*

#### `Error on ObjectSpecification #1: Unable to parse version portion of object reference nnnn/2/1, nnnn to an integer`&#x20;

UE: The wrong delimiter was used in the free text of the Assembly list or Genome list.&#x20;

*Change the delimiter to a semicolon (;) and try resubmitting the job.*

#### `Error on ObjectSpecification #1: Illegal number of separators /`

UE: The wrong delimiter was used in the free text of the Assembly list or Genome list.

*Change the delimiter to a semicolon (;) and try resubmitting the job.*&#x20;

#### `name assembly_info is not defined`&#x20;

UE: You are running the beta version of the app.&#x20;

*Change to the released version.*

## [Annotate Plant Transcripts with Metabolic Functions](https://narrative.kbase.us/#catalog/apps/kb_plant_rast/annotate_plant_transcripts/dbac462bc80e4d9db11efe0ff6e82c8fb28a3034)

#### `The genome does not contain any CDSs or features!`&#x20;

UE: The input genome does not have the needed features.&#x20;

*Double check the file and try resubmitting the job.*

## [**Assess Genome Quality with CheckM**](https://narrative.kbase.us/#catalog/apps/kb_Msuite/run_checkM_lineage_wf/release)

#### `Object NNN cannot be accessed: User user_name may not read workspace XXXXX` &#x20;

UE: The app requires data which is owned by another user and you do not have access. In another scenario, either you or the user have deleted the original assembly.&#x20;

*Ask the user for access to the needed file.*


# Functional Genomics App Errors

This is a list of error messages found within the Job Log for Functional Genomics App Jobs, what they mean and how to go about fixing them or if a job ticket needs to be submitted.

## **Types of Errors**

* **UE**=User fixable or possible user error
* **KE**=KBase error
* **UK**=Unknown or multiple possibles causes
* **TE**=Temporary system problem

{% hint style="info" %}
Use your Browser's search tool to paste in your error message to locate your error message, information about the error and next steps.&#x20;
{% endhint %}

You can always submit a ticket for help, questions, or follow-up to the [KBase Help Board](https://kbase-jira.atlassian.net/).&#x20;

## Cluster Expression Data: [Hierarchical](https://narrative.kbase.us/#catalog/apps/KBaseFeatureValues/expression_toolkit_cluster_hierarchical/), [Estimate K](https://narrative.kbase.us/#catalog/apps/KBaseFeatureValues/expression_toolkit_estimate_k/), [K-Means](https://narrative.kbase.us/#catalog/apps/KBaseFeatureValues/expression_toolkit_cluster_k_means/)

#### `Error running service CLI for method 'ClusterServiceR.cluster_hierarchical' with exit code 1 …...`

KE: The app is failing and needs maintenance.&#x20;

*Nothing can be done at this time. KBase is working on a solution to this problem.*&#x20;

## [BLASTp prot-prot Search](https://narrative.kbase.us/#catalog/apps/kb_blast/BLASTp_Search/)

#### `cannot have both input_one_sequence and input_one_ref parameter`&#x20;

UE: The app requires either a single Query Sequence or a Query Object with single sequence.&#x20;

*Change the inputs and try resubmitting the job.*

#### `output_one_name parameter required if input_one_sequence parameter is provided`&#x20;

UE: If a Query Sequence is used as input, an Output Query Sequence must be named.&#x20;

*Add an output name or change the input. Then try resubmitting the job.*

## [**Build a Feature Set**](https://narrative.kbase.us/#catalog/apps/FeatureSetUtils/build_feature_set/)

#### `Feature ID AT2G30640 does not exist in the supplied genome` &#x20;

UE The selected feature does not exist in the genome.&#x20;

*Recheck the selection and try resubmitting the job.*

## **Insert** [**Genome**](https://narrative.kbase.us/#catalog/apps/SpeciesTreeBuilder/insert_set_of_genomes_into_species_tree/) **or** [**Set of Genomes**](https://narrative.kbase.us/#catalog/apps/SpeciesTreeBuilder/insert_genomeset_into_species_tree) **into SpeciesTree**

#### `Error processing genome [xxx/xxx/x] (Not one protein family member found)`&#x20;

UE: The species tree is built using the predicted proteins in the genomes.&#x20;

*Run an app that does gene calling/genome annotation on the listed genome(s). Either RAST or Prokka annotation will work.*&#x20;

## [**Phylogenetic Pangenome Accumulation**](https://narrative.kbase.us/#catalog/apps/kb_phylogenomics/view_pan_phylo/)

#### `float division by zero`&#x20;

The app will fail with this message if there are only 2 genomes in the pangenome.&#x20;

*Try running the app with 3 or more genomes.*&#x20;

## **View Function Profile for** [**FeatureSet**](https://narrative.kbase.us/#catalog/apps/kb_phylogenomics/view_fxn_profile_featureSet/)**,** [**Genomes**](https://narrative.kbase.us/#catalog/apps/kb_phylogenomics/view_fxn_profile/)**, or** [**a Phylogenetic Tree**](https://narrative.kbase.us/#catalog/apps/kb_phylogenomics/view_fxn_profile_phylo/)

#### `ABORT: You must run the RAST SEED Annotation App or use SKIP option….`&#x20;

UE: The app is expecting RAST annotation for the input genome(s).&#x20;

*The easiest option is to check the 'Skip missing genomes' box in the advanced parameters. For more complete output 1) rune Annotate Microbial Genome(s) with RAST on the listed genomes and 2) use the genome(s) newly created by RAST as the input to the View Function Profile. Because the names are different, the app does not know how to find the newer version of the genome.*

#### `ABORT: You must run the 'Domain Annotation' App or use SKIP option ….`&#x20;

UE: The app is expecting Domain annotation for the input genome(s).

*Either run the Domain Annotation on the listed genome(s) or check the ‘Skip missing genomes’ box in the advanced parameters.*&#x20;

## [Classify Taxonomy of Metagenomic Reads with Kaiju](https://narrative.kbase.us/#catalog/apps/kb_kaiju/run_kaiju)

#### `missing or empty krona input file` &#x20;

UE: The filters were too restrictive and no output was generated.&#x20;

*Change the input parameters and try resubmitting the job.*

## [**GTDB-Tk classify**](https://narrative.kbase.us/#catalog/apps/kb_gtdbtk/run_kb_gtdbtk/release)

#### `(1, '/bin/bash -c "source activate py2 && GTDBTK_DATA_PATH` &#x20;

KE: Unknown error.&#x20;

*Nothing can be done at this time. KBase is working on a solution to this problem.*

#### `(2, '/bin/bash -c "source activate py2 && GTDBTK_DATA_PATH`&#x20;

UE: The ‘Minimum Alignment Percent’ was left blank&#x20;

*Assign a minimum percent in the advanced parameters and try resubmitting the job.*

## [**Create Differential Expression Matrix using DeSEQ2**](https://narrative.kbase.us/#catalog/apps/kb_deseq/run_DESeq2/release)&#x20;

#### `Invalid input:nselect Run All Paired Condition Combinations or provide partial condition pairs. Dont do both` &#x20;

UE: The user failed to select ALL vs Partial conditions. It cannot be blank and you cannot select both.&#x20;

*Do **one** of the following:*

1. *Check the box next to 'Run All Paired Condition Combinations.'*
2. *Add a 'Run Partial Condition Combinations.'*

## [Classify Taxonomy of Metagenomic Reads with GOTTCHA2](https://narrative.kbase.us/#catalog/apps/gottcha2/run_gottcha2/)

#### `(2, No such file or directory)`&#x20;

1. UE: The wrong reference database was used.&#x20;

   *Change the reference database and try resubmitting the job.*
2. KE: More than two reads libraries were included in the input.&#x20;

   *The developers have a task to fix this. The workaround is to only submit a maximum of two libraries at a time. Merging reads libraries can help with this process.*&#x20;

## [**Kraken**2 Taxonomic Sequence Classifier](https://narrative.kbase.us/#catalog/apps/kraken2/run_kraken2/beta)

#### `You must enter either an input genome or input reads`&#x20;

UE: There are two input options and one of them must be selected.&#x20;

*Add an input and try resubmitting the job.*

## [**VirSorter**](https://narrative.kbase.us/#catalog/apps/VirSorter/run_VirSorter/release)

#### `Cannot write data to fasta`&#x20;

KE: The user interface allows users to enter a genome as input but the app will reject it. The error has been reported.&#x20;

*Change to an input that is not a genome and try resubmitting the job.*


# Modeling App Errors

This is a list of error messages found within the Job Log for Modeling App Jobs, what they mean and how to go about fixing them or if a job ticket needs to be submitted.

## **Types of Errors**

* **UE**=User fixable or possible user error
* **KE**=KBase error

{% hint style="info" %}
Use your Browser's search tool to paste in your error message to locate your error message, information about the error and next steps.&#x20;
{% endhint %}

You can always submit a ticket for help, questions, or followup to the [KBase Help Board](https://kbase-jira.atlassian.net/jira/your-work).&#x20;

## [**Bulk Download Modeling Objects**](https://narrative.kbase.us/#catalog/apps/fba_tools/bulk_download_modeling_objects/release)

#### `Must provide one and only one of workspace name`&#x20;

UE: No modeling objects were found the the user’s narrative.&#x20;

*Add some objects and try resubmitting the job.*

#### `GLPK does not support ID's longer than 256 characters`&#x20;

KE: This option fails when this condition exists and the app needs maintenance.

*Nothing can be done at this time. KBase is working on a solution to this problem.*&#x20;

## [**View Flux Network**](https://narrative.kbase.us/#catalog/apps/fba_tools/view_flux_network/)

#### `Can't use an undefined value as an ARRAY reference`&#x20;

KE: The app is failing and needs maintenance.&#x20;

*Nothing can be done at this time. KBase is working on a solution to this problem.*&#x20;

#### `Authentication required for AbstractHandle but no authentication header was passed`&#x20;

KE: The app is failing and needs maintenance.&#x20;

*Nothing can be done at this time. KBase is working on a solution to this problem.*&#x20;

## [**Compare Models**](https://narrative.kbase.us/#catalog/apps/fba_tools/compare_models/)

#### `Must select at least two models to compare`&#x20;

UE: The app requires two or more models for the comparison.&#x20;

*Add at least one more model and and try resubmitting the job.*


# Developing Apps

Interested in developing an App for KBase? Here you can find information for developers.

{% hint style="danger" %}

## **Due to infrastructure changes and upgrades, we are not currently accepting new developer applications at this time.**

{% endhint %}

## **The Principles**

### **Reproducibility**

KBase is a scientific platform. A cornerstone of science is that scientific experiments and analyses are reproducible - that is a scientist must be able to take a description of an experiment and/or analysis, perform said experiment / analysis independently, and get the same results. Science that is not reproducible is not science, and knowledge that cannot be verified independently is not knowledge.

### **Provenance**

Provenance in the context of KBase explains how data comes to exist - the sequence of operations that transformed a set of units of data into a different set of units of data along with who caused those transformations to occur and when. Information on job performance and the error log should be considered part of the data produced. Accurate provenance enables many of the other KBase principles, including reproducibility and Credit where credit is due. Data without provenance is not useful, as there is no way to determine how the data was created and therefore assess the data’s reliability.

### **Sharing and data privacy**

Users loading data into and creating data within KBase are guaranteed that their data and activities are private unless they explicitly share their data or make it public. KBase will not mine, collate, or otherwise use their private data or activities/jobs (other than for internal tasks required to administer the platform), and their data/activities/jobs will not be visible in any data view to users that are not granted access.

When data or analyses are shared, they are expected to be viewable and runnable by the users they are shared with. If a user makes a copy of data or analyses, those data and analyses are expected to be viewable and runnable just as the sources are viewable and runnable. Information about the original generators of the data are expected to propagated along with the data to ensure proper credit.&#x20;

### **Credit where credit is due**

Users that add data, apps, or analyses to the system and make them available for other users to rerun, reuse, or copy and modify must receive credit for their work.

### **Support of cross-references**

Data in KBase should reference appropriate and relevant related data e.g., a Genome object referencing the Taxon object to which it belongs. Data should be richly connected to other data types when feasible. Additionally, data imported into the system should reference the data source.

### **FAIR data management**

[**Findable, Accessible, Interoperable, Reusable**](https://www.nature.com/articles/sdata201618) data meets the criteria set out in the linked paper. FAIR overlaps with many of the other principles set out here but is an emerging standard for data management. Further, we wish to ensure data of the same type are comparable i.e., data are in comparable units derived from similar “workflows” so for example, numbers from two different data sets can be fairly combined.&#x20;

### **Open source software and data**

KBase software is publicly available through [GitHub](https://github.com/kbase). You may reuse KBase code under the terms of our [open source license](https://github.com/kbase/sdkbase2/blob/master/LICENSE.md). All software and data used in KBase must be open source and restriction free. This, along with reproducibility and provenance allows users to examine and understand each step in analyses workflows.

### **Extensible by 3rd parties**

3rd parties can extend the functionality of KBase by contributing application modules that run in the KBase execution environment. These modules are expected to also follow the principles outlined here.

## [Develop and release your own App](https://www.kbase.us/develop/)

If you want to develop Apps using the SDK, please apply for a KBase developer account by going to the Requesting a [New KBase Developer Account](https://accounts.kbase.us/index.php?tpl=request_identity.tpl). If you are a US citizen, your account can be created within a few days. For foreign nationals, it may take several weeks (and, in a few cases, you may not be able to get a developer account). Non-US citizens will be asked for additional information in order to process their application.Once your account is approved, [contact us](https://www.kbase.us/support/) with your username and ask to be added to the developer list.

### KBase System Architecture

KBase is based loosely on a service-oriented architecture that bundles related functionality into a set of independently scalable services that are managed to provide responsive interaction via the Narrative Interface.

For more details, please see the KBase [architecture overview](https://github.com/kbase/KBaseDeveloperBootstrap/blob/master/README.md), which outlines the components and relationships between KBase’s user interfaces, services and databases.


# The KBase SDK

KBase Software Development Kit for collaborators

{% hint style="danger" %}

## Due to infrastructure changes and upgrades, we are not currently accepting new developer applications at this time.

{% endhint %}

## Information for Developers

The [KBase Software Development Kit (SDK)](https://kbase.github.io/kb_sdk_docs/) offers members of the KBase community a mechanism to add open-source, open-license (as [defined by OSI](https://opensource.org/licenses)) analysis tools to KBase so that they run on [KBase’s computational architecture](https://github.com/kbase/KBaseDeveloperBootstrap/blob/master/README.md) and are available through the [KBase Narrative Interface](https://narrative.kbase.us/). Community developers can write new tools that take advantage of the [KBase Structured Data Types](https://narrative.kbase.us/#catalog/datatypes) and large[ KBase Reference Data Resources](https://www.kbase.us/data-policy-and-sources/). Check which tools have already been made available in our [App Catalog](https://kbase.us/applist/) and add your own!

## Get Started

### 1. [Sign up for a user account](https://narrative.kbase.us/#signup)

You need a user account to be able to access KBase’s user interfaces, APIs, workspaces, and more.

### 2. [Create a KBase developer account](/development/create-a-kbase-developer-account)

### 3. Read the [SDK Documentation](https://kbase.github.io/kb_sdk_docs/)

Learn to use [KBase SDK](https://kbase.github.io/kb_sdk_docs/) to begin crafting your tools to add to the KBase [App Catalog](https://kbase.us/applist/).

{% embed url="<https://youtu.be/Q6qt7gqaVnM>" %}


# Create a KBase Developer Account

{% hint style="danger" %}

## **Due to infrastructure changes and upgrades, we are not currently accepting new developer applications at this time.**

{% endhint %}

The following instructions outline how to create a [KBase Developer Account](http://accounts.cels.anl.gov/). This has two phases 1) apply for an Argonne collaborator account and 2) join the KBase-Developers project.&#x20;

### Apply for an Argonne Collaborator Account.&#x20;

1. Go to the Computing, Environment and Life Sciences Account Management Create a Collaborator Account Page: [https://apps.anl.gov/registration/collaborators](< https://apps.anl.gov/registration/collaborators>).&#x20;
2. Read Terms and then click 'Accept Terms and Conditions' button.\
   ![](https://lh5.googleusercontent.com/D_5FSgEQSnCUEtnmbs7k5t-LuY8tOrYTY3x4vK2Vu06LXa9PjMBEVcOLTU21HDUZ8HO4x9pThbLWxvD7vXbZMGyVYg-yvZbWD9rkiMCyRoiIsVxbIqYIR8IpVon7FPS-fW1pnC54RQrfF63MYX-VU8NcDmOuELbXEy6eZOB58SjyMGfEr4CRj77IS7eMx5m-YUNhWTwfwA)
3. Fill out sponsor form with the following information: \
   *Sponsor E-mail*: <chenry@anl.gov>\
   *Reason for Account*: KBase developer account for \[your KBase dev username here]\
   Then click 'Verify Sponsor E-mail'.\
   ![](/files/IL7HO0abSysZvb9ph5gB)
4. Fill out the Argonne Collaborator Registration form. \
   *Please note: Make sure to set “Are you a US citizen” correctly. If you are not, you will be required to fill out additional information and provide full documentation.*\
   ![](https://lh4.googleusercontent.com/7htIgaaGceW4DZsfNZvMXWxhUd8O6HaL0Mu6b1LROTv1fTsgLwTv0EwsOBWvpleL6GhRe_JIrJBHWxLeR57yzxlA2Ee4kOivjYN6YqGk1Iz4J4AYn7UydZ0IbkdJIgSgyS87v9toexgdom-_vZ6-VmXrtRrv3bgu7y0Gd-WkoxwkAeg1O-PnvORQYGpVP_427muaFLG7LQ)
5. Click the Create Collaborator button when complete.&#x20;
6. **Email&#x20;*****<engage@kbase.us>*****&#x20;to let us know that you have requested a collaborator account!**&#x20;
7. Once your account has been approved, you will receive an email to set your password. Follow the instructions to *change* your initial password. \
   ![](https://lh5.googleusercontent.com/wa5MwwzBDhDKwpZX2YW6rvtbJz2jhUIJa2vZKjPXO7qJPgYqrKshow7Wc7NLhMj4gBkXIYDs6-fY-Bu3DcWIXlYcQt5u8TtRnC-u5gk3PkIpmn3yN_dgIVZMmNXHdGzWRNAG1_7LfiGBglem6Yw5Pcwwwl6hmqH7RE5T03A9KJ-0rf24zL-agLZmpLFwz9SxUNuQuuxY9g)
8. You can now login to [https://accounts.cels.anl.gov](<https://accounts.cels.anl.gov >). \
   ![](https://lh6.googleusercontent.com/4moAjPiLXI97bZxJVJ0N8b_HKjlGxL1ll-f4fmHTkehVH7ARf__W8Wn7ICk0z-clzV8xFtquBfi7p2_Y04UgN1YTrscnrdhe0Ro_lLdlM7xghIUf9TObCoWLqI84GYr8cghT-THXCdNiq3QWG6C5kl8sJ7TchYuIJ8gE-92rEsKgDde-RB1cf24e0IIkMyqyjMeo9k5LQg)

### Join the KBase-Developers Project

1. Login to [https://accounts.cels.anl.gov](<https://accounts.cels.anl.gov >).&#x20;
2. To join the KBase-Developers Group, click on 'Join Project' in the left menu bar. \
   ![](https://lh3.googleusercontent.com/HqzJpF98WnCpiX5CGXdHVS-Nf-Hs5NRUSLl-gehzj55jQ21vNCK5skwTNssO3QyzrJDHGuYMp4bhlrI49v5BsIV1NvAo-BL18N3jWDvnAwePqhth4N0RwUuGe-lHEwkgtc2WHZvvLcaIfFMPCHdFdAkmeXW-iSMSXg6sYILgAAGyhF0_dutO6YXSPTwVlfSyoYqTPurNAQ)
3. Search for "KBase-Developers" and click on 'KBase-Developers'.\
   ![](https://lh5.googleusercontent.com/ZYVCeXUiB3C5lghDsa5n_ebIbxzLW76Sv59IqxXO3aa3ybU2McPd4X4ia02zR43UUFiY9QgPW27MOMEAn08cX59c7KsuuzWRx2nWD34LDpsqK9UvmvIZKi5OliRc95YGAXYEqUrMfwqVs02PyrxUpVO8tpicbJRqh8nDLLV-cefs1lrHLo05xpiOgHw3OA12Tz1oCo0eTA)
4. Click on the 'Request Membership' button.\
   ![](https://lh4.googleusercontent.com/ElVrHoD6BDWRzhqvqDpu70_4Soc9hXX1DVqduyjP_MxukImtSnsHYSYqjbbpaJ6nd9s0rocdkSXpBId_2dUuaCapdehSYpqQGDmr70XRmRHjd3D84GS_rTFJDvOGmSAN62CIEfnFDFEIk-jBnwweIy_JArEiMR4lr-sepKjEf_6RLaH5LqySmafLG1qfrsVdJtglDzsWMw)
5. **Using the same thread as started above, email&#x20;*****<engage@kbase.us>*****&#x20;to let us know that you have requested to join the KBase-Developers Group along with your KBase developer username.**

{% hint style="warning" %}
Note - KBase Developer accounts must be renewed on a yearly basis. You will receive a notification email from KBase staff regarding your renewal. If your account expires prior to renewal, please register with the SAME USERNAME, and we will do our best to expedite your permissions to the same account.
{% endhint %}


# KBase GitHub Repository

## Access the [KBase GitHub Repository here](https://github.com/kbase).

For more details, see the KBase [architecture overview](https://github.com/kbase/KBaseDeveloperBootstrap/blob/master/README.md), which outlines the components and relationships between KBase’s user interfaces, services and databases.


# kbase.us

Specific how to edit and use constraints in WordPress.

## Responsive Image Sizing

For webpages viewed on a larger screen i.e., computer, the maximum width is 800px.&#x20;

If image width is greater than 800px

`width="100%” height="auto”`

If width is less than 800px

`max-width="100%” height="auto”`

If width is not 800px or greater, but you want to set a specific maximum size.&#x20;

`width="100%” max-width=”NNNpx” height="auto”`

## Add padding to an image

After the size, add the following, specifying side of image and space:

```
style="padding-left: Npx" 
style="padding-right: Npx"
```

## Creating boxes around text

`<p style="border: Npx; border-style:solid; border-color: #HTMLtag; padding: 1em;"> text here </p>`

Teal boxes for developer highlights are 5px and #009688.&#x20;

## Anchors

## Stop text from wrapping around an align left image

\<a id="anchor">\</a>

```
<br clear="left">
```

Add anchor to end of slug for web address

<https://siteaddress/slug/#anchor>

## Embedding and Linking from Drive

1. Open the PDF or image file in a new preview window
2. Open the expanded menu and click on "Embed item..."
   1. To embed an image, select the link or frame.
   2. To link a PDF, copy the Drive link.&#x20;
3. Insert the link in desired location.&#x20;

## Pop up messages

Best practices for using Hustle for messaging on kbase.us

1. Add or edit descriptive Title including dates.
2. Include informative Main Content.&#x20;
3. Add enable the "Never see this message again" link
4. Optional, schedule specific dates, but revert to "Draft" after maintenance period.&#x20;
5. Publish!


