Skip to content

Authentication and authorisation - #145

Open
khaledk2 wants to merge 45 commits into
ome:mainfrom
khaledk2:authentication_and_authorisation
Open

khaledk2 wants to merge 45 commits into
ome:mainfrom
khaledk2:authentication_and_authorisation

Conversation

@khaledk2

Copy link
Copy Markdown
Collaborator

Overview

This PR introduces authentication and authorisation to the search engine for serving omero data sources. It also introduces the concept of public and private data sources.

For public data sources, there is no change to how the search engine is used. For private data sources, users must obtain a JWT token and include it with each request.

Authentication

A JWT token can be obtained through the /auth/login endpoint. The user must provide:

  • Username
  • Password
  • Data source name

Once the searcher receives these credentials, it authenticates the user with the OMERO Server. If authentication is successful, the search engine returns a JWT token. The token contains the user's authorisation information, including their user ID and groups. This information is encoded in the token and must be included with subsequent requests to the search engine.

The default token lifetime is two hours. This is configurable, and an administrator can modify the expiration time using the set_JWT_expire_time method in the command.py script.

The JWT token must be included in the request headers. A dedicated example script, auth_token_query.py, has been added to the examples directory. It demonstrates how to:

  1. Authenticate and obtain a JWT token.
  2. Include the token in subsequent search requests.

Swagger API Documentation

Authorisation has also been added to the Swagger API documentation.

To use the authenticated APIs through Swagger:

  1. Obtain a JWT token using the /auth/login endpoint.
  2. Click the Authorise button on the right-hand side of the Swagger UI.
  3. Enter the token.
  4. Use the private data source APIs with the authenticated session.

Testing

I have tested the implementation with a local omero-server and compared the search results from:

  • The search engine directly.
  • A Python script connected to the omero-server.

The results were consistent across these different methods.

However, more extensive testing is still required to cover different authentication, authorisation, and data-source scenarios.

case_sensitive,
bookmark,
pagination_dict,
raw_elasticsearch_query,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should raw_elasticsearch_query be removed from this call now (unused)?

for group in owned_groups:
print(f"Name: {group.getName()} | ID: {group.getId()}")
if group.getId() not in groups:
groups[group.getId()] = {"name": group.getName()}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think that all the groups where you are an owner will be included in the conn.getGroupsMemberOf() so you shouldn't find any extra ones in listOwnedGroups().

E.g. if I print conn.getEventContext() I see:

    memberOfGroups = 
    {
        [0] = 53
        [1] = 1
        [2] = 3
        [3] = 1575
    }
    leaderOfGroups = 
    {
        [0] = 1575
        [1] = 3
    }

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The check for private groups is currently commented out for debugging purposes. I will uncomment it to ensure that private groups are not added during the first loop. The second loop will then add the group(s) if the user is the owner.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

After that change, I won't see any private groups that I am a member of, unless I am also an owner of that group? Why would we want that behaviour (which is different from what the OMERO permissions system allows).

@khaledk2 khaledk2 Sep 14, 2026 •

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think in an OMERO private group, only two types of users (plus server administrators) can view data inside a private group:

  • The Owner of the data
  • The Group Owner

The results will be based on both the user and their associated groups. This permission is applied within this method.:
get_permission_query insided omero_search_engine/api/v1/resources/utils.py

Comment thread omero_search_engine/__init__.py Outdated
or (
token
and token.get(data_source)
and not token.get(data_source).get("is_valid")

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It's a little tricky to understand the logic here.
E.g if token.get(data_source) is false, and token.get("is_valid") is true then you won't get the "private datasource" error below. Maybe that will "never" happen, or you'll get some other error later, rather than a security violation.
But it might be clearer to do:

if not token:
  return "Error, no token"
if not token.get("is_valid"):
  return "Error: invalid token"
if not token.get(data_source):
  return "Error: no data source"
if not token.get(data_source).get("is_valid"):
  return "Error: data source is private"

@will-moore

Copy link
Copy Markdown
Member

I'm not clear from the description of how access to "private" data sources is controlled.
Is there a separate data source for every group in OMERO, and I am allowed to access a data source for every group that I am a member of?
If that is true, when I list /data_sources/ it might be nice to know which data source contains data from which group?

Is it possible for unauthenticated users to be able to access the list of private data_sources that they don't have access to? If so then I think we should not allow that since even the names of the data sources may contain confidential information?

@khaledk2

Copy link
Copy Markdown
Collaborator Author

I'm not clear from the description of how access to "private" data sources is controlled. Is there a separate data source for every group in OMERO, and I am allowed to access a data source for every group that I am a member of? If that is true, when I list /data_sources/ it might be nice to know which data source contains data from which group?

Is it possible for unauthenticated users to be able to access the list of private data_sources that they don't have access to? If so then I think we should not allow that since even the names of the data sources may contain confidential information?

Each data source has a PUBLIC flag that indicates whether the data source instanc is public or private. A value of true means the data source is public, while false means it is private.

For example, IDR should have:

PUBLIC: true

For the private data source, e.g. night shade, it should be:
PUBLIC: false

I have restricted the visibility of the data source name for unauthenticated users.

@will-moore

Copy link
Copy Markdown
Member

For the private data sources, how does the user (admin) control what is indexed?
Does all the data in "group A" get indexed into "datasource A" and all the data from "group B" get indexed into "datasource B"?
Then the members of group A can search "datasource A" but not "datasource B"?
And if the user gets added to group B then they can immediately search "datasource B"?

The definition of what is "public" data in OMERO.web is simply what the public user can access.
If the public user gets added to a group, then the data in that group becomes public, and if removed from a group then it becomes "private". This is unlikely to happen often, but we still need to consider this scenario (may need to re-index when this happens)?

@khaledk2

Copy link
Copy Markdown
Collaborator Author

The data source represents where the data comes from and is independent of whether the data is public or private. A data source can therefore contain data that is subject to different access permissions.

The user's ID and group memberships are obtained during login and are then used to determine which data the user can access when performing a search.

For private data, access is controlled through group membership. The user's permissions determine which data they can search, rather than restricting access to an entire data source.

For example, if a user is a member of Group A, they can search the data that Group A has permission to access. If they are also added to Group B, they will then be able to search the additional data that Group B has permission to access.
Similarly, if a user is removed from a group, they should lose access to the data associated with that group's permissions.

The definition of "public" in OMERO.web is based on what an unauthenticated/public user can access. If a public user is added to a group, they would gain access to the data permitted by that group. If they are subsequently removed from the group, that data would no longer be accessible to them.

Because these permissions are evaluated when performing the search, changes to group membership should not require the data to be re-indexed. The search permissions can be evaluated against the user's current group membership at query time.

@will-moore

will-moore commented Sep 15, 2026 •

Copy link
Copy Markdown
Member

So, when the data sources are built, how do you determine which data is indexed in which datasource?
Is there any difference between "public" and "private" data sources when they are created?

The OMERO server itself has no concept of "public" and "private" data, but the PR description above "[This PR] introduces the concept of public and private data sources" sounds like there is a distinction between public and private datasources.

But then your last comment above sounds like there's no difference between public and private datasources since it is all determined by which groups the "public user" is in when they perform the search.
The "public user" is purely defined by the config in OMERO-web, so how does the searchengine know who is the public user?

If a datasource contains a mixture of data from different groups, then the search itself must add extra clauses to only retrieve data that a particular user is able to access? E.g. something like: ...AND group is in (3, 4)?

@khaledk2

Copy link
Copy Markdown
Collaborator Author

Yes, there is a flag on the indexed data which indicates the data source it belongs to.
The public/private distinction is a property of the data source configuration. Public and private data sources are otherwise similar, with the main difference being the public flag.
A public data source means that all data inside that data source has public access, regardless of the data owner or the groups associated with the data. Public data sources do not require access control. For example, a public data source can contain data indexed from CSV files, where authentication is not currently involved.
This public data source concept is important so that existing implementations, such as the IDR Gallery, can continue to work with this branch without requiring any code changes.
The public user will be configured as part of the data source configuration, and this concept will be implemented in the search engine so that it can determine the appropriate access when performing a search when no user is logged in to a private data source.
For private data sources, access control is applied based on the user's access rights. If a data source contains data from different groups, the search engine adds the appropriate clauses to the query based on the groups/data that the user is authorised to access. Conceptually, this could be something like:

... AND group IN (3, 4)

So, in summary:

  • Public data source: all data in the data source is public, regardless of owner or groups; no access control is required.
  • Private data source: access control is applied based on the user's access rights.
  • Search query: for access-controlled data, the appropriate access clauses are added based on the user's permissions/groups.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants