---
sourceDocument: Australia ServiceNow AI Platform Administration
sourceDocumentLink: https://servicenow-prod.fluidtopics.net/r/platform-administration

 Release :

    - australia

ft:locale :

    - en-US

ft:publication_title :

    - Australia ServiceNow AI Platform Administration

ft:clusterId :

    - platadm

bundleId :

    - platadm

workflow :

    - Platform


---

# Configure crawl settings

# Configure crawl settings for a GitLab external content connector {#ariaid-title1}

Release version: Australia  
Updated May 15, 2026  
![](https://www.servicenow.com/docs/portal-asset/ico-clock) 6 minutes to read  
Specify the groups, projects, and repositories you want your GitLab external content connector to crawl. Select the issues, wikis, merge requests, tags, branches, and commits you want the crawl to retrieve and feed to AI Search for indexing.

## Before you begin

A connector administrator must have already created the GitLab external content connector that you want to configure crawl settings for. To learn about this procedure, see [Create a GitLab external content connector](https://servicenow-prod.fluidtopics.net/hBkxEqspGOa1xG2ZrhXQpQ "Create an external content connector to retrieve searchable content and security principals from your GitLab source system.").

Role required: sn_ext_conn.xcc_admin

## About this task

This task is optional. By default, the GitLab external content connector crawls content from all subgroups, projects, and repositories found in top-level groups owned by the GitLab.com user that it's configured to impersonate, and sends all supported content types (issues, wikis, merge requests, tags, branches, and commits) to AI Search for indexing. Only perform this task if you want the connector to use any of the following non-default settings:

* Inclusion or exclusion filters for the subgroups to crawl when running content crawls
* Inclusion or exclusion filters for the projects/repositories to crawl when running content crawls
* Inclusion or exclusion filters for the types of content to retrieve from the source system when running content crawls
* Inclusion or exclusion filters for the branches to retrieve from the source system when running content crawls
{#configure-crawl-settings-gitlab-external-content-connector__ul_onp_x1f_tdc}

Content is only retrieved from the source system if it passes all of your configured crawl setting filters. If any crawl setting filter excludes a
content item, the external content connector doesn't retrieve it.  
Important:  
By default, each external content connector can index up to one million (1,000,000) content items from its source system. When a connector exceeds this limit, it continues to crawl the source system,
but only sends content item deletions and updates to AI Search for indexing, ignoring new content items. The connector logs an error message for every 10,000 content items it crawls beyond the indexing limit.{#configure-crawl-settings-gitlab-external-content-connector__store-app-external-content-connectors-indexing-limit-numeric-ph}

When a connector's indexed content item count exceeds 800,000, a warning message appears in the connector's UI to indicate that it's approaching the indexing
limit. If the connector reaches the indexing limit, an error message appears in its UI.

External content connectors that support user permissions crawls can handle permissions for up to five hundred thousand (500,000) users and their groups. If a connector retrieves
users in excess of this limit, user and group permissions may not be correctly applied to the connector's retrieved content. As a result, the content may not be searchable.

If one of your connectors reaches the content indexing limit, you can update its crawl
settings and file inclusion/exclusion filters to reduce the number of content items it retrieves. Alternatively, if you need a connector to index more than 1,000,000 content items, you can create a Customer Service and Support case at <https://support.servicenow.com/now> to request a limit increase for the connector.

## Procedure

1. Navigate to AllExternal Content ConnectorsExternal Content Admin Home. {#configure-crawl-settings-gitlab-external-content-connector__navigate-ext-content-connectors-admin-home-step}
{#configure-crawl-settings-gitlab-external-content-connector__navigate-ext-content-connectors-admin-home-step}
2. In the Connectors list, select the record for the GitLab external content connector whose settings you want to modify.
3. In the connector editor's Settings tab, select Crawl settings. {#configure-crawl-settings-gitlab-external-content-connector__select-crawl-settings-step}
{#configure-crawl-settings-gitlab-external-content-connector__select-crawl-settings-step}
4. Select one of the following Group filtering options:
   * To crawl all subgroups found in top-level groups owned by the connector's impersonated GitLab.com user account, select Crawl all groups.
   * To crawl only a specified set of subgroups found in top-level groups owned by the connector's impersonated GitLab.com user account, select Include only these groups, then use the Add group URLs to include field and Add button to
     enter URLs for the groups that you want to include in the crawl.

     For example, you might enter <kbd class="ph userinput">https://gitlab.com/example-dot-com/production</kbd> to include only searchable content from the <kbd class="ph userinput">production</kbd> subgroup and all subgroups that it
     contains.
   * To crawl all except a specified set of groups found in top-level groups owned by the connector's impersonated GitLab.com user account, select Exclude only these groups, then use the Add group URLs to exclude field and Add button to
     enter URLs for the groups that you want to exclude from the crawl.

     For example, you might enter <kbd class="ph userinput">https://gitlab.com/example-dot-com/test-*</kbd> to exclude searchable content from all subgroups with names that start with <kbd class="ph userinput">test-</kbd>.

   {#configure-crawl-settings-gitlab-external-content-connector__choices_cdr_s4x_sdc}  
   Note:  
   Subgroup inclusion URLs can be specified as prefixes, with the wildcard character <kbd class="ph userinput">*</kbd> at the end of the URL matching any string.
5. Select one of the following Project/repository filtering options:
   * To crawl all projects and repositories owned by the connector's impersonated GitLab.com user account, select Crawl all projects/repositories.
   * To crawl only a specified set of projects and repositories owned by the connector's impersonated GitLab.com user account, select Include only these projects/repositories, then use the Add project/repository URLs to include field and Add button to enter URLs for the projects and repositories that you want to include in the crawl.  
     Note:  
     Project and repository inclusion URLs can be specified as prefixes, with the wildcard character <kbd class="ph userinput">*</kbd> at the end of the URL matching any string.

     For example, you might enter <kbd class="ph userinput">https://gitlab.com/example-dot-com/prod-*</kbd> to include only searchable content from projects whose names start with <kbd class="ph userinput">prod-</kbd>.
   * To crawl all except a specified set of projects and repositories owned by the connector's impersonated GitLab.com user account, select Exclude only these projects/repositories, then use the Add project/repository URLs to exclude field and Add button to enter URLs for the projects and repositories that you want to exclude from the crawl.  
     Note:  
     Project and repository exclusion URLs can be specified as prefixes, with the wildcard character <kbd class="ph userinput">*</kbd> at the end of the URL matching any string.

     For example, you might enter <kbd class="ph userinput">https://gitlab.com/example-dot-com/confidential273</kbd> to exclude searchable content from the <kbd class="ph userinput">confidential273</kbd> project.
   {#configure-crawl-settings-gitlab-external-content-connector__choices_thl_q2b_4fc}
6. Enable the Crawl content types options for the types of content you want to retrieve when you run content crawls.  
   The GitLab external content connector supports indexing of searchable content for these content types:{#configure-crawl-settings-gitlab-external-content-connector__table_lw5_fjk_tfc__entry__2}

   | Content type | Searchable content indexed |
   |-|-|
   | Issues | Issue description |
   | Wikis | MarkDown content converted to HTML (without attachments) |
   | Merge requests | Merge request description (MarkDown) and discussions |
   | Tags | Tag message |
   | Branches | Commit message of head commit |
   | Commits | Commit message |
   [ ]

   {#configure-crawl-settings-gitlab-external-content-connector__table_lw5_fjk_tfc}  
   Important:  
   The GitLab external content connector doesn't support indexing of searchable content from any of these content types:
   * Commit, issue, and wiki discussions
   * Commit diffs
   * Content from archived groups or projects
   * Content from groups or projects in the pending deletion state
   * Content from subgroups of top-level groups that aren't owned by the impersonated GitLab.com user
   * Content of files attached to issues or merge requests
   * Content of wiki attachments in formats other than plain text (.txt)
   * Internal or confidential notes in merge request discussions
   * Repository files
   {#configure-crawl-settings-gitlab-external-content-connector__ul_kgm_pjk_tfc} {#configure-crawl-settings-gitlab-external-content-connector__select-content-types-step}
{#configure-crawl-settings-gitlab-external-content-connector__select-content-types-step}
7. If you included the Branches content type in step [6](https://servicenow-prod.fluidtopics.net/bvIJuSgwIAATSfrY~s9WCA#configure-crawl-settings-gitlab-external-content-connector__select-content-types-step), use the Add branches to include in regex format field and the Add button to specify Java regular expression patterns matching the names of branches you want to include in content crawls.  
   As an example, you might specify <kbd class="ph userinput">^2025.*$</kbd> to include branches with names that start with <kbd class="ph userinput">2025</kbd>, or specify <kbd class="ph userinput">^.*$</kbd> to crawl all branches. To learn about Java regular expression pattern syntax, see [the Javadoc for the java.regex.util.Pattern class](https://docs.oracle.com/en/java/javase/21/docs/api/java.base/java/util/regex/Pattern.html).  
   Note:  
   The branch name expressions \^main$ and \^master$ are included by default. You can't remove these branches from the list.
8. **Optional:** If you want AI Search to automatically generate captions for content in attachments and files retrieved by the connector, select the Multimodal captions option.  
   When you select this option, the Platform Multimodal Service automatically generates descriptive captions for images, tables, charts, and other visual elements found in retrieved attachments and files. You can discover these retrieved attachments and files by searching for terms from their generated captions.  
   This option is only available when the Platform Multimodal Service plugin is activated on your instance.
   * For details on activating the plugin, see [Activate the Platform Multimodal Service plugin](https://servicenow-prod.fluidtopics.net/Aps2A3ZsA_gtaZTILcqu2w "Enable automatic generation of searchable descriptive captions for images, tables, charts, and other visual elements found in indexed attachments.").
   * To learn how to select the VLM (visual learning model) provider and model used for the Platform Multimodal Service, see [Configure multimodal captioning for AI Search](https://servicenow-prod.fluidtopics.net/nE2_exg8skwGj6J7f2cuwQ "Use the AI Search Admin console to select the visual language model (VLM) provider and model for multimodal captioning.").
9. Select Save and validate.
{#configure-crawl-settings-gitlab-external-content-connector__steps_yqk_jjw_22c}

## Result

The GitLab external content connector is updated with your modified crawl settings.

## What to do next

To retrieve content from your GitLab source system using your modified crawl settings, create and run a one-time content crawl for your GitLab external content connector. To learn about creating and running one-time content crawls, see [Create a content crawl for an external content connector](https://servicenow-prod.fluidtopics.net/fFJ92ZYif96fkc2JA4HtyQ "Retrieve searchable content and metadata from your source system with a content crawl. Run the crawl as a one-time task or schedule it to run on a recurring basis.").

*[\>]: and then


