Conversation
…horisation for private data sources
for more information, see https://pre-commit.ci
| case_sensitive, | ||
| bookmark, | ||
| pagination_dict, | ||
| raw_elasticsearch_query, |
There was a problem hiding this comment.
Should raw_elasticsearch_query be removed from this call now (unused)?
| for group in owned_groups: | ||
| print(f"Name: {group.getName()} | ID: {group.getId()}") | ||
| if group.getId() not in groups: | ||
| groups[group.getId()] = {"name": group.getName()} |
There was a problem hiding this comment.
I think that all the groups where you are an owner will be included in the conn.getGroupsMemberOf() so you shouldn't find any extra ones in listOwnedGroups().
E.g. if I print conn.getEventContext() I see:
memberOfGroups =
{
[0] = 53
[1] = 1
[2] = 3
[3] = 1575
}
leaderOfGroups =
{
[0] = 1575
[1] = 3
}
There was a problem hiding this comment.
The check for private groups is currently commented out for debugging purposes. I will uncomment it to ensure that private groups are not added during the first loop. The second loop will then add the group(s) if the user is the owner.
There was a problem hiding this comment.
After that change, I won't see any private groups that I am a member of, unless I am also an owner of that group? Why would we want that behaviour (which is different from what the OMERO permissions system allows).
There was a problem hiding this comment.
I think in an OMERO private group, only two types of users (plus server administrators) can view data inside a private group:
- The Owner of the data
- The Group Owner
The results will be based on both the user and their associated groups. This permission is applied within this method.:
get_permission_query insided omero_search_engine/api/v1/resources/utils.py
| or ( | ||
| token | ||
| and token.get(data_source) | ||
| and not token.get(data_source).get("is_valid") |
There was a problem hiding this comment.
It's a little tricky to understand the logic here.
E.g if token.get(data_source) is false, and token.get("is_valid") is true then you won't get the "private datasource" error below. Maybe that will "never" happen, or you'll get some other error later, rather than a security violation.
But it might be clearer to do:
if not token:
return "Error, no token"
if not token.get("is_valid"):
return "Error: invalid token"
if not token.get(data_source):
return "Error: no data source"
if not token.get(data_source).get("is_valid"):
return "Error: data source is private"
|
I'm not clear from the description of how access to "private" data sources is controlled. Is it possible for unauthenticated users to be able to access the list of private data_sources that they don't have access to? If so then I think we should not allow that since even the names of the data sources may contain confidential information? |
Each data source has a For example,
For the private data source, e.g. I have restricted the visibility of the data source name for unauthenticated users. |
|
For the private data sources, how does the user (admin) control what is indexed? The definition of what is "public" data in OMERO.web is simply what the public user can access. |
|
The data source represents where the data comes from and is independent of whether the data is public or private. A data source can therefore contain data that is subject to different access permissions. The user's ID and group memberships are obtained during login and are then used to determine which data the user can access when performing a search. For private data, access is controlled through group membership. The user's permissions determine which data they can search, rather than restricting access to an entire data source. For example, if a user is a member of Group A, they can search the data that Group A has permission to access. If they are also added to Group B, they will then be able to search the additional data that Group B has permission to access. The definition of "public" in OMERO.web is based on what an unauthenticated/public user can access. If a public user is added to a group, they would gain access to the data permitted by that group. If they are subsequently removed from the group, that data would no longer be accessible to them. Because these permissions are evaluated when performing the search, changes to group membership should not require the data to be re-indexed. The search permissions can be evaluated against the user's current group membership at query time. |
|
So, when the data sources are built, how do you determine which data is indexed in which datasource? The OMERO server itself has no concept of "public" and "private" data, but the PR description above "[This PR] introduces the concept of public and private data sources" sounds like there is a distinction between public and private datasources. But then your last comment above sounds like there's no difference between public and private datasources since it is all determined by which groups the "public user" is in when they perform the search. If a datasource contains a mixture of data from different groups, then the search itself must add extra clauses to only retrieve data that a particular user is able to access? E.g. something like: |
|
Yes, there is a flag on the indexed data which indicates the data source it belongs to.
So, in summary:
|
for more information, see https://pre-commit.ci
…/khaledk2/omero_search_engine into authentication_and_authorisation
Overview
This PR introduces authentication and authorisation to the search engine for serving omero data sources. It also introduces the concept of public and private data sources.
For public data sources, there is no change to how the search engine is used. For private data sources, users must obtain a JWT token and include it with each request.
Authentication
A JWT token can be obtained through the
/auth/loginendpoint. The user must provide:Once the searcher receives these credentials, it authenticates the user with the OMERO Server. If authentication is successful, the search engine returns a JWT token. The token contains the user's authorisation information, including their user ID and groups. This information is encoded in the token and must be included with subsequent requests to the search engine.
The default token lifetime is two hours. This is configurable, and an administrator can modify the expiration time using the
set_JWT_expire_timemethod in thecommand.pyscript.The JWT token must be included in the request headers. A dedicated example script,
auth_token_query.py, has been added to theexamplesdirectory. It demonstrates how to:Swagger API Documentation
Authorisation has also been added to the Swagger API documentation.
To use the authenticated APIs through Swagger:
/auth/loginendpoint.Testing
I have tested the implementation with a local omero-server and compared the search results from:
The results were consistent across these different methods.
However, more extensive testing is still required to cover different authentication, authorisation, and data-source scenarios.